Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:33:55.284758Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2606.29296.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:33:55.284758Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4bd02a22-1010-4a61-a1af-0e9881eb6641 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1cdd8a9e-171f-450b-9f4f-b7a2e0a06c6c · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2104129-86f3-4840-ba2f-df6f312ef891 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61182028-d306-4652-8739-f1b2efda3b50 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5141abb-d49c-4149-a2f3-10b441144e9d · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9121599e-67e2-4b46-b717-87f431398145 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Fipo: Eliciting deep reasoning with future-kl influenced policy optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2089c9ec-877d-485d-a7de-9a276d833b30 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Junxi Yin, Haisen Luo, Zhenyu Li, Yihua Liu, Dan Liu, Zequn Li, and Xiaohang Xu
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0a1ea47-ae02-4361-a8f3-bb93f2e628c7 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Pinpointing crucial steps: Attribution-based credit assignment for verifiable reinforcement learning.arXiv preprint arXiv:2510.08899
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c43312f5-f21f-47e1-9ba6-50aae585c62f · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9a15bad-183c-43fa-90d7-579b34f4bb28 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Drpo: Efficient reasoning via decoupled reward policy optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 164d8e26-45ef-4cf3-8df2-4949f61a554c · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 21ccff55-542f-47bb-8b8e-6372587e1828 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c210f49-f540-4104-a341-0d2b6abc61a3 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6311e93f-6bf9-4425-b939-c5455ad5fd4c · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f91f9df-6b75-48ce-a74c-fe82f04854bc · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af4dd9d7-5f6a-49ad-963d-d50f7f2a66b5 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners ♫ M u S i Q ue: Multihop Questions via Single-hop Question Composition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 98ac186d-3007-4402-a640-ff7c727d445f · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Qwen2.5 Technical Report
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35ecaa5c-aadf-4b27-a957-aff9c423f466 · outbound
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Setting k<1.0 protects necessary reasoning verbosity and raises the exploratory ceiling; k=0.7 achieves the best average pass@1 and is used as the default throughout this paper
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.