Pith. sign in

Paper Citation Record · LEDGER

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2506.00027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00027 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:32.681228Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:56:43.392662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:56:44.691483Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bba346f9-7c64-4b17-b385-303de0e8fc09 · outbound

This paper cites This loss function inherently accommo- dates the probabilistic nature of the labels generated via Monte Carlo simulations.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This loss function inherently accommo- dates the probabilistic nature of the labels generated via Monte Carlo simulations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.465025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:31.706585Z digest=sha256:6018e9d4f505d7a8a39cefa54fe42a41b275e04d7448a681464eb1021848c937

Observation 7c16b16c-e274-4a28-b43d-960efeba1e36 · outbound

This paper cites Key hy- perparameters such as learning rate, batch size, and weight decay are tuned based on validation perfor- mance.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Key hy- perparameters such as learning rate, batch size, and weight decay are tuned based on validation perfor- mance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.355153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:31.800088Z digest=sha256:3ef49de2079a626b6763ef3aaaa35d8d8cdca06277d5dc11d80fc9f60f7916c9

Observation 8d42cbdf-8536-4493-ab1b-cd6705899bb5 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:35.196714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:31.912875Z digest=sha256:ba38d612ab43eb3b52da1d6b68b6a17f648f1d3f8fd523535a79e9c2cca551a2

Observation 44e147b2-6a4d-4e99-8d78-306d6385c705 · outbound

This paper cites Through this comprehensive training process, the PRM learns to accurately predict the correct- ness of intermediate steps in reasoning tasks.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Through this comprehensive training process, the PRM learns to accurately predict the correct- ness of intermediate steps in reasoning tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:35.092280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.011816Z digest=sha256:70680833ea7e88e789aaf9854e90ca63ece4888b727c44a34ce4abe5604b9712

Observation 5f56dcf1-dcf5-464f-9bd0-3ece17db7a02 · outbound

This paper cites Generation: Generate N candidate solutions for a given problem.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate N candidate solutions for a given problem

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.970889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.085667Z digest=sha256:ae009e8466e8a80d32ee60db525063ebd4819aaba8143332f5e983ee138c78cc

Observation cc9e8e8a-119b-4968-930e-78aa750a4660 · outbound

This paper cites Initialization: Start with an initial set of paths (beam width K).

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Initialization: Start with an initial set of paths (beam width K)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.838749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.134622Z digest=sha256:4385363174e601a6ab87ce94113ddfe2db3632a5eddb0a4303d0a56a9999840d

Observation ee66bcbe-b6fb-4051-a049-d1597e554db1 · outbound

This paper cites Tree Structure: Represent the reasoning process as a tree where nodes are states and edges are actions.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Tree Structure: Represent the reasoning process as a tree where nodes are states and edges are actions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.684994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.169968Z digest=sha256:dae136fdbabe8bed85662fbc73814251aafb4d7728e3991fb67bae16ca916c68

Observation 85045de8-8db1-461e-a55c-14529b4d9488 · outbound

This paper cites Generation: Generate mul- tiple candidate solutions.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate mul- tiple candidate solutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.562022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.237742Z digest=sha256:315b01a98a7f633eb28c41dbf3fa4f1ceee47fa8c61ce9333413eef9af2b01de

Observation d1516759-e0b1-48dc-90de-691ecb764514 · outbound

This paper cites Evaluate x1 using the PRM to obtain a correctness score p1.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate x1 using the PRM to obtain a correctness score p1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.427544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.294762Z digest=sha256:481c11d079ebd3bf615c5e697c1579f70eff5021a10b2422156a81d51707f614

Observation 76db70f9-f352-43b1-8125-5f83a158852d · outbound

This paper cites Evaluate each candidate step using the PRM, selecting the one with the highest score.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate each candidate step using the PRM, selecting the one with the highest score

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.312072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.319712Z digest=sha256:fb3cf0558c5af49e10776a3c44b787a0d5f45418af0f66c9a407cc3e0dd13529

Observation b637c00a-7475-43e2-ad36-7bab6de445ea · outbound

This paper cites Use strategies such as PRM-Last (considering the score of the last step) or PRM-Min (considering the minimum score across steps) to determine the final solution.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Use strategies such as PRM-Last (considering the score of the last step) or PRM-Min (considering the minimum score across steps) to determine the final solution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.145604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.359064Z digest=sha256:76e57dfe8d62a3168148c94f8f49141e4c5ef73e6c9404f20c0a5f5c9fe284dc

Observation ce9a8883-6a9a-42ad-856f-8deb5b74f96d · outbound

This paper cites Utilize modern hardware ac- celerators, such as GPUs, to handle the increased computational load.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Utilize modern hardware ac- celerators, such as GPUs, to handle the increased computational load

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:34.019206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.398936Z digest=sha256:54d753a8ff8e613cf265b0239bd76b01a22f76af41983c4639f97c340b9fc3b6

Observation ab19f1b9-ac38-4461-87e1-1ace9ebd8450 · outbound

This paper cites Imple- ment adaptive strategies to balance the depth and breadth of search, optimizing resource usage.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Imple- ment adaptive strategies to balance the depth and breadth of search, optimizing resource usage

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.823336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.439421Z digest=sha256:a45a8b0c37ce9657b81cd9e0fa1b13f20b369665766c8ea5060e3e3953de5ac6

Observation 4654c7ba-005c-402b-b5d3-1040160d8ae1 · outbound

This paper cites Prune low-scoring paths early in the search process to focus computational efforts on promising candidates.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Prune low-scoring paths early in the search process to focus computational efforts on promising candidates

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.662415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.495362Z digest=sha256:a121b77a0386fa083707a66090806934dc22a1e94d0490c581ea519162622d77

Observation cc8ebed9-c271-4d82-819f-98e0a79f680b · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.465675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.534588Z digest=sha256:390498555374d1e7fe2067e892a515953b658e246bf0f8ec23263a1a6834bc2b

Observation 0248f2c1-d0f9-4fb8-86aa-459924e9acc5 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.229148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.572618Z digest=sha256:b599d96eb2e9351fc64b6848dafd62038b6945b15fa05a8fdb97c6f0cb064689

Observation 8a0d53e8-e69c-49f7-9e74-deaea23ec7e7 · outbound

This paper cites This adaptability is par- ticularly valuable in real-world applications where resource constraints may vary.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This adaptability is par- ticularly valuable in real-world applications where resource constraints may vary

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:33.119005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.632868Z digest=sha256:e71d6876f90273437d66c4e6040435ce74d63129d209ffc2441fb5cd645f0949

Observation 22f9940c-0462-4b55-a407-3765bdd45aa0 · outbound

This paper cites an unresolved cited work.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:31:33.013658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:32.681228Z digest=sha256:15c6e9e24ecee3fa315c18dbcd9028fac3eb629337ce820bc6f8c77b53901a00

Observation 2c0c9326-0eed-4aff-9ab8-f92e75dac285 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.378407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.378407Z digest=sha256:03898bf11ae9e0e1c10912f8c86a3841a309170d4ba91233c02a003c9be0930d

Observation 448e7d3d-9561-438e-9162-830cae32bb52 · outbound

This paper cites Let's Verify Step by Step.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.448994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.448994Z digest=sha256:4a451a2aa8bcfa4d8a66fe12c20fb426d625cbe25b3c96533c7ac55b310a2309

Observation 0c39dbd5-9b15-40a9-9920-fcc822bf7ff5 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.553358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.553358Z digest=sha256:f5e975d954ab02412962d75fd80042fb4401aa9c224d92eef7eecd8faff42a39

Observation d5444646-d5de-4dac-9b70-5200418c738f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:31.630290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:31.630290Z digest=sha256:b4c2d7f80f80ce49a982717c5f2e02a7f822cb90a7f391389772f562e06ac66d

Pith citing papers

Observation 1d3a22a1-49e8-45c1-85e7-701a2f4f4e21 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:56:44.698316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:56:43.392662Z digest=sha256:9a1ac4024a9c5e8d7d667d9d7e76d2fa1f00041ec940542607dbc579509826a1

Observation 9abfc3ba-89cc-483d-b9d9-d37af6e6f29c · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:76e4a0dbb785009295aad2f24c75429ecea0492d12426f8e1aefed97d9547d45