Pith. sign in

Paper Citation Record · LEDGER

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2506.00027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00027 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:32.681228Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:56:43.392662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:56:44.691483Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bba346f9-7c64-4b17-b385-303de0e8fc09 · outbound

This paper cites This loss function inherently accommo- dates the probabilistic nature of the labels generated via Monte Carlo simulations.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This loss function inherently accommo- dates the probabilistic nature of the labels generated via Monte Carlo simulations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.465025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:31.706585Z digest=sha256:df6ba1b7e75086ddea73ac54148489edd1877513b16ad58bed493dc0f662285f

Observation 7c16b16c-e274-4a28-b43d-960efeba1e36 · outbound

This paper cites Key hy- perparameters such as learning rate, batch size, and weight decay are tuned based on validation perfor- mance.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Key hy- perparameters such as learning rate, batch size, and weight decay are tuned based on validation perfor- mance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.355153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:31.800088Z digest=sha256:e4896ff9e156ae995f05cf54b8879b6bf11b3f51e00db25b00745466210e9695

Observation 8d42cbdf-8536-4493-ab1b-cd6705899bb5 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:35.196714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:31.912875Z digest=sha256:1daf0d3fab3fcbc8f4d7e321bd1cf65aa6bd846a523f655454332a706fb558ab

Observation 44e147b2-6a4d-4e99-8d78-306d6385c705 · outbound

This paper cites Through this comprehensive training process, the PRM learns to accurately predict the correct- ness of intermediate steps in reasoning tasks.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Through this comprehensive training process, the PRM learns to accurately predict the correct- ness of intermediate steps in reasoning tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.092280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.011816Z digest=sha256:4e76a360dfd0150383e6fddd58c3e38d6d1806a3c700f3aa6901987cf92a6dce

Observation 5f56dcf1-dcf5-464f-9bd0-3ece17db7a02 · outbound

This paper cites Generation: Generate N candidate solutions for a given problem.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate N candidate solutions for a given problem

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.970889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.085667Z digest=sha256:3a0c4eb316fb5ff086a8c021344dc64a006b344d27ff55c60b7dad7d1a5651eb

Observation cc9e8e8a-119b-4968-930e-78aa750a4660 · outbound

This paper cites Initialization: Start with an initial set of paths (beam width K).

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Initialization: Start with an initial set of paths (beam width K)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.838749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.134622Z digest=sha256:7b21d5f037c906ca7c893b77abf116d9e83bf6848e11c12de3945ba05cea2838

Observation ee66bcbe-b6fb-4051-a049-d1597e554db1 · outbound

This paper cites Tree Structure: Represent the reasoning process as a tree where nodes are states and edges are actions.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Tree Structure: Represent the reasoning process as a tree where nodes are states and edges are actions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.684994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.169968Z digest=sha256:2e501595ba0004d5f3485e8dd4de709c84cb7d40af7fc38e581effbb3d79b835

Observation 85045de8-8db1-461e-a55c-14529b4d9488 · outbound

This paper cites Generation: Generate mul- tiple candidate solutions.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate mul- tiple candidate solutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.562022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.237742Z digest=sha256:05406b91e9506e57e770eb459679e7ee00d080755ccef551b97f11df3aa11e72

Observation d1516759-e0b1-48dc-90de-691ecb764514 · outbound

This paper cites Evaluate x1 using the PRM to obtain a correctness score p1.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate x1 using the PRM to obtain a correctness score p1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.427544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.294762Z digest=sha256:85608f8804aa33e7a6e803c5c584ffac1eb5a0cc6fbd8cd57e33a8def86f7365

Observation 76db70f9-f352-43b1-8125-5f83a158852d · outbound

This paper cites Evaluate each candidate step using the PRM, selecting the one with the highest score.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate each candidate step using the PRM, selecting the one with the highest score

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.312072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.319712Z digest=sha256:4d840df50c5ded458652f62e4e3299ffb8502b08c2e57f03e39d3d52037470d5

Observation b637c00a-7475-43e2-ad36-7bab6de445ea · outbound

This paper cites Use strategies such as PRM-Last (considering the score of the last step) or PRM-Min (considering the minimum score across steps) to determine the final solution.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Use strategies such as PRM-Last (considering the score of the last step) or PRM-Min (considering the minimum score across steps) to determine the final solution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.145604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.359064Z digest=sha256:68dbe6c5747ef9360eaed6676d9012dc574c131d4ed392f5f119f5fed3e7b493

Observation ce9a8883-6a9a-42ad-856f-8deb5b74f96d · outbound

This paper cites Utilize modern hardware ac- celerators, such as GPUs, to handle the increased computational load.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Utilize modern hardware ac- celerators, such as GPUs, to handle the increased computational load

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.019206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.398936Z digest=sha256:5f4f3171ed14eb14e53a2230a82d76c1cb979ed8b44eb1506fa02a9cbe741fe7

Observation ab19f1b9-ac38-4461-87e1-1ace9ebd8450 · outbound

This paper cites Imple- ment adaptive strategies to balance the depth and breadth of search, optimizing resource usage.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Imple- ment adaptive strategies to balance the depth and breadth of search, optimizing resource usage

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.823336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.439421Z digest=sha256:20443eaa4e0d3a0a6d4d71f49de6d5336fd1fb34c03ea8810bc5436449c8799a

Observation 4654c7ba-005c-402b-b5d3-1040160d8ae1 · outbound

This paper cites Prune low-scoring paths early in the search process to focus computational efforts on promising candidates.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Prune low-scoring paths early in the search process to focus computational efforts on promising candidates

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.662415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.495362Z digest=sha256:f5dd722930b7130b851a744d6cc1c0a3e011186b8614bb1d6328d006bbdb418d

Observation cc8ebed9-c271-4d82-819f-98e0a79f680b · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.465675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.534588Z digest=sha256:21f09e5f3e346c03eb3daa83a3653e10428c3d7596a2fbcd9052931e24c71216

Observation 0248f2c1-d0f9-4fb8-86aa-459924e9acc5 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.229148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.572618Z digest=sha256:a05c2269a3f4dab1929c93553b70030ecf9620d25295b7dc1498945e199ab74e

Observation 8a0d53e8-e69c-49f7-9e74-deaea23ec7e7 · outbound

This paper cites This adaptability is par- ticularly valuable in real-world applications where resource constraints may vary.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This adaptability is par- ticularly valuable in real-world applications where resource constraints may vary

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.119005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.632868Z digest=sha256:8536f4cdef32eb293642912a82c489b7c39ac1fd35de774b5c2460103a39528c

Observation 22f9940c-0462-4b55-a407-3765bdd45aa0 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.013658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:32.681228Z digest=sha256:a4b0c4be41db6c423a18947836be4974f6354bc39687dc020aa2c2d895d54495

Observation 2c0c9326-0eed-4aff-9ab8-f92e75dac285 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.378407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.378407Z digest=sha256:03898bf11ae9e0e1c10912f8c86a3841a309170d4ba91233c02a003c9be0930d

Observation 448e7d3d-9561-438e-9162-830cae32bb52 · outbound

This paper cites Let's Verify Step by Step.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.448994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.448994Z digest=sha256:4a451a2aa8bcfa4d8a66fe12c20fb426d625cbe25b3c96533c7ac55b310a2309

Observation 0c39dbd5-9b15-40a9-9920-fcc822bf7ff5 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.553358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.553358Z digest=sha256:f5e975d954ab02412962d75fd80042fb4401aa9c224d92eef7eecd8faff42a39

Observation d5444646-d5de-4dac-9b70-5200418c738f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.630290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.630290Z digest=sha256:b4c2d7f80f80ce49a982717c5f2e02a7f822cb90a7f391389772f562e06ac66d

Pith citing papers

Observation 1d3a22a1-49e8-45c1-85e7-701a2f4f4e21 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:56:44.698316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T17:56:43.392662Z digest=sha256:200cffd3e5656d6f4d2f673b3b8680bacbadb697a38f0de6e3acd6d79b399f72

Observation 9abfc3ba-89cc-483d-b9d9-d37af6e6f29c · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:76e4a0dbb785009295aad2f24c75429ecea0492d12426f8e1aefed97d9547d45