Pith. sign in

Paper Citation Record · LEDGER

Truncated Proximal Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2506.15050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15050 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:13:59.500195Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.662953Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a07db9a-f440-4f6d-bc78-de05eddd17e3 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Truncated Proximal Policy Optimization

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.459940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:d48c005c4393278bff0a95012f0ce536d57ee047328009e74f1e1dbd67a77f2b

Observation 74d765c9-4b30-47bc-90f7-46531bec998a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Truncated Proximal Policy Optimization

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.344183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:7b0b7193ee79e42316bb8cd4bf63b47a6381c177f11c509c50da556fde3d8a7d

Observation 07d8bc75-86d9-4f0a-9f9f-4e93c609926e · inbound

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning cites this paper.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning Truncated Proximal Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.136801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:29.657563Z digest=sha256:4523f90aec13350333aa4d8bb51e256c2c39afac56fc396c330d7f7a512304de

Observation 724cda01-6204-4d9f-a62b-2f24ba64b398 · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models Truncated Proximal Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.820470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:efba9a97cbcf9b508937c5a0ff7ec5a5aed2cfd7cb225ab4c5d9ba75f25ce4a2

Observation 6b352224-e2f1-4ad2-95fa-4ea7464b7a21 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Truncated Proximal Policy Optimization

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.702861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:79527d08f1aa30cfb73347ec19d62071fd79ea127f2230d17df4cc432ef2e22d

Observation 75abb5a4-9f43-4b1d-be3f-db03143579a6 · inbound

Escaping the KL Agreement Trap in On-Policy Distillation cites this paper.

Escaping the KL Agreement Trap in On-Policy Distillation Truncated Proximal Policy Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:29.494171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T16:58:57.106367Z digest=sha256:c4e2ba86eaac55592aa4fb252df6d6036c3cda98ddda2969b73fa04e3b2e09ff

Observation 5ce08037-3f17-43ef-96a0-4eaddc27b022 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Truncated Proximal Policy Optimization

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.664267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:33cd9681a75279c6dba90041baca177037d2b0048a94d2beca271ed2a8b934b2

Observation 3bb90ee1-5fd7-4296-b7f4-225c94d6c918 · inbound

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF cites this paper.

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF Truncated Proximal Policy Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:58.800569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T01:23:07.441218Z digest=sha256:25b84bde7f6473427144e74d115638a1ec4399f632d760dd6d044788f2590737

Observation 64ffcc2a-3cd9-4eea-bcba-de093c193ae0 · inbound

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training cites this paper.

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training Truncated Proximal Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:13:59.500195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:13:59.500195Z digest=sha256:0e899e64075058fa61e7a2da7af93bd43b4b18ef2ca580f19cb36bf8cb954234