Pith. sign in

Paper Citation Record · LEDGER

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2606.29296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29296 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:33:55.284758Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bd02a22-1010-4a61-a1af-0e9881eb6641 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.180773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:a98fdabd45ff95660999c4ea68546024f3e424cd244a0a9472b625c98d5fa5d6

Observation 1cdd8a9e-171f-450b-9f4f-b7a2e0a06c6c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:34:20.679568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:c2e77e1df1409c6a10f12caf9a893a682454465e405014fff42f8da31b142be8

Observation a2104129-86f3-4840-ba2f-df6f312ef891 · outbound

This paper cites Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.178202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:8710a2306a765c12562d9a6084dde3839f707476d23c73e8fc86ff36b0e9e0d8

Observation 61182028-d306-4652-8739-f1b2efda3b50 · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.175163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:d7b05545376e7012506c86dfceaa54096efda94e452d7e9cecda11e4f9fa9afa

Observation b5141abb-d49c-4149-a2f3-10b441144e9d · outbound

This paper cites Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.201958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:96cde3b401486614bcab93334e018c91018142b979b20a5775b82cfc81c93059

Observation 9121599e-67e2-4b46-b717-87f431398145 · outbound

This paper cites Fipo: Eliciting deep reasoning with future-kl influenced policy optimization.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Fipo: Eliciting deep reasoning with future-kl influenced policy optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:34:21.183878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:4868e161bfe5e6b0a352c7aa16159c50ca566e539797917abfa1979ac31a30cd

Observation 2089c9ec-877d-485d-a7de-9a276d833b30 · outbound

This paper cites Junxi Yin, Haisen Luo, Zhenyu Li, Yihua Liu, Dan Liu, Zequn Li, and Xiaohang Xu.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Junxi Yin, Haisen Luo, Zhenyu Li, Yihua Liu, Dan Liu, Zequn Li, and Xiaohang Xu

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.208021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:ab51da226305389e3883b052375d1102e50010caa0e3236bda18304027444b5f

Observation d0a1ea47-ae02-4361-a8f3-bb93f2e628c7 · outbound

This paper cites Pinpointing crucial steps: Attribution-based credit assignment for verifiable reinforcement learning.arXiv preprint arXiv:2510.08899.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Pinpointing crucial steps: Attribution-based credit assignment for verifiable reinforcement learning.arXiv preprint arXiv:2510.08899

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:34:21.198548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:ec8bb25f65db812b0ad4e5aa0ed18d6296d82d4eee6f89cd279f9800c78af0a3

Observation c43312f5-f21f-47e1-9ba6-50aae585c62f · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 9

Resolution
malformed identifier
doi_truncated, observed 2026-06-30T07:34:20.679855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:d376d28e821333c8baca16fc56e7ef12027b592cde8e1ad5b24d4a9773a4dbce

Observation d9a15bad-183c-43fa-90d7-579b34f4bb28 · outbound

This paper cites Drpo: Efficient reasoning via decoupled reward policy optimization.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Drpo: Efficient reasoning via decoupled reward policy optimization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:34:21.205151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:dd61c4e95823b251450c18d99560793f535920d48535d5642f127eb4a405e995

Observation 164d8e26-45ef-4cf3-8df2-4949f61a554c · outbound

This paper cites Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:39:30.770950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:230a46db5f0372b6c8cfd5f3374c424113732772ed47b870dd7af3c51f6058cd

Observation 21ccff55-542f-47bb-8b8e-6372587e1828 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:34:21.186840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:30c57c4431722aaa875a265c76ba1ad7f7ba5a6bef707e043a68b4b98fc22210

Observation 5c210f49-f540-4104-a341-0d2b6abc61a3 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.192165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:e6e1f87145c994c353de777d08270c1190a3eb92adef2b2e23088ae3e7f89046

Observation 6311e93f-6bf9-4425-b939-c5455ad5fd4c · outbound

This paper cites an unresolved cited work.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-30T07:33:55.284758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:432b7796931e26828b2043b882180540cb727917fcecdd48977aa92a6c55bdfb

Observation 0f91f9df-6b75-48ce-a74c-fe82f04854bc · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 15

Resolution
metadata mismatch
doi, observed 2026-06-30T07:34:20.683703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:c8aed192d24496a26f70f39cb1fa7ed9644ca8ecd509fb9baf9383c58a4fb079

Observation af4dd9d7-5f6a-49ad-963d-d50f7f2a66b5 · outbound

This paper cites ♫ M u S i Q ue: Multihop Questions via Single-hop Question Composition.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners ♫ M u S i Q ue: Multihop Questions via Single-hop Question Composition

Reference 16

Resolution
metadata mismatch
doi, observed 2026-06-30T07:34:20.689750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:9f147ab90087177c1aecbebae4999566a045e9b7b195601f11585fcff97d8cd0

Observation 98ac186d-3007-4402-a640-ff7c727d445f · outbound

This paper cites Qwen2.5 Technical Report.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Qwen2.5 Technical Report

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:34:21.189518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:bd5eccaad8abd50d1c6ac3925bc785f35e952428cf5c90a6d1c4bd6d0c40d507

Observation 35ecaa5c-aadf-4b27-a957-aff9c423f466 · outbound

This paper cites Setting k<1.0 protects necessary reasoning verbosity and raises the exploratory ceiling; k=0.7 achieves the best average pass@1 and is used as the default throughout this paper.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Setting k<1.0 protects necessary reasoning verbosity and raises the exploratory ceiling; k=0.7 achieves the best average pass@1 and is used as the default throughout this paper

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-30T07:33:55.284758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:e4e93c68edf980886b4f1995e59402e521fd1ece6449afdd76a638e4ba057106

Pith citing papers

No inbound Pith citation observations are available.