Pith. sign in

Paper Citation Record · LEDGER

Truncated Proximal Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2506.15050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15050 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:13:59.500195Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.662953Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a07db9a-f440-4f6d-bc78-de05eddd17e3 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Truncated Proximal Policy Optimization

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.459940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:919e28d78829ca1f764f70b97e9ec46d30432084c5dd90ec58541a7ddd2574e3

Observation 74d765c9-4b30-47bc-90f7-46531bec998a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Truncated Proximal Policy Optimization

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.344183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:b455ec3042c0531f1ba89eb7699fedbeebd166eb4e978fadebb3a73fe5ad3466

Observation 07d8bc75-86d9-4f0a-9f9f-4e93c609926e · inbound

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning cites this paper.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning Truncated Proximal Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.136801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:09:29.657563Z digest=sha256:fac8a666d5926789232cd842fb35c93c9ea09c0b66ee42f37b7a838fc90da56d

Observation 724cda01-6204-4d9f-a62b-2f24ba64b398 · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models Truncated Proximal Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.820470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:eecd16403aa00e0b198016fb583980240adb24b6fc8988ce749e61126d09765d

Observation 6b352224-e2f1-4ad2-95fa-4ea7464b7a21 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Truncated Proximal Policy Optimization

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.702861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:ecab692f7ab74b76832090f2f5eefa341d2cd6e6a733868a20b0089dc0b94fa5

Observation 75abb5a4-9f43-4b1d-be3f-db03143579a6 · inbound

Escaping the KL Agreement Trap in On-Policy Distillation cites this paper.

Escaping the KL Agreement Trap in On-Policy Distillation Truncated Proximal Policy Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:29.494171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:58:57.106367Z digest=sha256:7a80cf2256e5595aa8d366e98184884e3231e5540c36e7350a6a7565c48f1299

Observation 5ce08037-3f17-43ef-96a0-4eaddc27b022 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Truncated Proximal Policy Optimization

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.664267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:6157d29208e4a37161d111472331a65c8924bbaa2827204fcf5878375bdd8d84

Observation 3bb90ee1-5fd7-4296-b7f4-225c94d6c918 · inbound

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF cites this paper.

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF Truncated Proximal Policy Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:58.800569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T01:23:07.441218Z digest=sha256:389332fb17ad0dc20b676e46a0e3b0207fc40231563a5b712f554a393e5e39cb

Observation 64ffcc2a-3cd9-4eea-bcba-de093c193ae0 · inbound

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training cites this paper.

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training Truncated Proximal Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:13:59.500195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:13:59.500195Z digest=sha256:0e899e64075058fa61e7a2da7af93bd43b4b18ef2ca580f19cb36bf8cb954234