Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T23:24:43.556606Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2507.15698.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T23:24:43.556606Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T11:06:21.527926Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T11:54:38.713100Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5b2b89f6-d0dd-4ee0-b31e-116eee59ea4d · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Qwen Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 163fec65-6cc6-462f-b789-08baac871e75 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning ODIN: Disentangled Reward Mitigates Hacking in RLHF
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6de869e4-ab8d-4720-967e-6ff994714573 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning The Llama 3 Herd of Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d88245e-cb38-489a-962f-89e09abc6976 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0457bb7-171b-43e5-998b-d3f28515fca2 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 832db875-eaa1-4760-875a-c1ea9ec09a22 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6dfc41c7-39f8-4962-a5ff-d7e3e45a9360 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Chen et al
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1e3dcf4-64f9-4e84-bc91-83c1113b8d8d · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1dff3c38-1b04-47eb-a071-2298d56aea64 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Process Reward Model with Q-Value Rankings
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 024e02c0-bee5-4045-8315-4b0e7c3c85a7 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Let's Verify Step by Step
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed1b5cf5-91ac-4b26-9453-d01605e7cf36 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning DeepSeek-V3 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d27e7ee-a36b-4ce5-9891-32da91502027 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9adddc11-31b6-4e55-a087-ba9a64a6eb06 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning LLM Critics Help Catch LLM Bugs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65ee2097-0693-4d23-9a6e-91d17f2a279f · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d9cd287-48e2-4169-96c8-fb4ab79daf80 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62c27199-f8e2-4e43-9735-7098902b0d17 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3d41963-09c7-4cd6-a177-9b69d02f6db5 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46083a4f-feed-45d9-bf47-7d29457ad3d1 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 24e63f3c-ad87-44a4-b51f-3942077965e3 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fedc1ac9-4c04-441a-bf70-3ec5c4844a97 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86982e4b-f595-4084-a8d2-944086c00017 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e484e6dd-d2f0-4ee9-a0c9-fcdc4b3c3939 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 099e8384-6906-4034-833d-cc9237fa9108 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Evaluating Mathematical Reasoning Beyond Accuracy
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c9865bc-fc99-4cf6-956d-e6d378c1c85c · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Qwen3 Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 965e8885-c52a-4ce9-961b-875a5e9e0d61 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e0c21f7-4a31-4d49-aa00-3063fb45931f · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b93e8691-1a74-4db4-ae01-bb89c48410e0 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ebc0511-b75b-49f4-8847-cc68dee08fd6 · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b0b7362-6f40-4600-af13-74dd7cca7d5c · outbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3e94bf0-9133-432f-add4-393ededb352e · inbound
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 533848bc-7e3c-49a6-9437-236c1e7ff9b4 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95d3a8f1-771d-47c0-ae0c-9673fba8a254 · inbound
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.