Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:32.681228Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2506.00027.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:32.681228Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:56:43.392662Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:56:44.691483Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bba346f9-7c64-4b17-b385-303de0e8fc09 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This loss function inherently accommo- dates the probabilistic nature of the labels generated via Monte Carlo simulations
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c16b16c-e274-4a28-b43d-960efeba1e36 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Key hy- perparameters such as learning rate, batch size, and weight decay are tuned based on validation perfor- mance
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8d42cbdf-8536-4493-ab1b-cd6705899bb5 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 44e147b2-6a4d-4e99-8d78-306d6385c705 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Through this comprehensive training process, the PRM learns to accurately predict the correct- ness of intermediate steps in reasoning tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f56dcf1-dcf5-464f-9bd0-3ece17db7a02 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate N candidate solutions for a given problem
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc9e8e8a-119b-4968-930e-78aa750a4660 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Initialization: Start with an initial set of paths (beam width K)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee66bcbe-b6fb-4051-a049-d1597e554db1 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Tree Structure: Represent the reasoning process as a tree where nodes are states and edges are actions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85045de8-8db1-461e-a55c-14529b4d9488 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Generation: Generate mul- tiple candidate solutions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d1516759-e0b1-48dc-90de-691ecb764514 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate x1 using the PRM to obtain a correctness score p1
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 76db70f9-f352-43b1-8125-5f83a158852d · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Evaluate each candidate step using the PRM, selecting the one with the highest score
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b637c00a-7475-43e2-ad36-7bab6de445ea · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Use strategies such as PRM-Last (considering the score of the last step) or PRM-Min (considering the minimum score across steps) to determine the final solution
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce9a8883-6a9a-42ad-856f-8deb5b74f96d · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Utilize modern hardware ac- celerators, such as GPUs, to handle the increased computational load
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ab19f1b9-ac38-4461-87e1-1ace9ebd8450 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Imple- ment adaptive strategies to balance the depth and breadth of search, optimizing resource usage
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4654c7ba-005c-402b-b5d3-1040160d8ae1 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Prune low-scoring paths early in the search process to focus computational efforts on promising candidates
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc8ebed9-c271-4d82-819f-98e0a79f680b · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0248f2c1-d0f9-4fb8-86aa-459924e9acc5 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a0d53e8-e69c-49f7-9e74-deaea23ec7e7 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling This adaptability is par- ticularly valuable in real-world applications where resource constraints may vary
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22f9940c-0462-4b55-a407-3765bdd45aa0 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c0c9326-0eed-4aff-9ab8-f92e75dac285 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Training Verifiers to Solve Math Word Problems
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448e7d3d-9561-438e-9162-830cae32bb52 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Let's Verify Step by Step
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c39dbd5-9b15-40a9-9920-fcc822bf7ff5 · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5444646-d5de-4dac-9b70-5200418c738f · outbound
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3a22a1-49e8-45c1-85e7-701a2f4f4e21 · inbound
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9abfc3ba-89cc-483d-b9d9-d37af6e6f29c · inbound
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.