Pith. sign in

Paper Citation Record · LEDGER

Central Path Proximal Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2506.00700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00700 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:07:09.566879Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:41.630698Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.078046Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 074502be-9dda-46de-81e9-5d3b8d32d5fd · outbound

This paper cites A safe exploration approach to constrained Markov decision processes.

Central Path Proximal Policy Optimization A safe exploration approach to constrained Markov decision processes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.495244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:08.931227Z digest=sha256:3def744fb832019568e77c652852e74d09cd2995f9012945a73ad7eda246f8bf

Observation 78af1034-67c6-4a9f-8fd3-24376f09bebb · outbound

This paper cites Direct Behavior Specification via Constrained Reinforcement Learning.

Central Path Proximal Policy Optimization Direct Behavior Specification via Constrained Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.187120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:09.093230Z digest=sha256:40666196c0caa1923008fd5779f8e7592af309fc5bd1b38663d1dff35fa56902

Observation 4faecdac-4fa9-411c-a493-11ed459e1bb1 · outbound

This paper cites Responsive Safety in Reinforcement Learning by PID Lagrangian Methods.

Central Path Proximal Policy Optimization Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:09.904564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:09.280854Z digest=sha256:757849fea0924c27b9d7abb720a8446f9a22f8b27ab7237ad947de49112bdb51

Observation d55bf51f-2cc4-4f0f-8774-b0fda9159b60 · outbound

This paper cites Constrained Reinforcement Learning with Smoothed Log Barrier Function.

Central Path Proximal Policy Optimization Constrained Reinforcement Learning with Smoothed Log Barrier Function

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.385071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.385071Z digest=sha256:b33e977c9c88edc621aa58c2a5a7b0de0eeba10715fde47139847f156d9e0c0c

Observation d3628220-eedf-49dc-a3a3-71f7824ecd61 · outbound

This paper cites Yiming Zhang, Quan Vuong, and Keith Ross.

Central Path Proximal Policy Optimization Yiming Zhang, Quan Vuong, and Keith Ross

Reference 13

Resolution
verified exact
doi, observed 2026-08-07T12:07:09.751998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:09.480957Z digest=sha256:29be5e52d4398bae9820cfe00319f8ffb148fa146f520f0bf98448b3b17cdad0

Observation 5d638274-83d0-44c6-8aef-e4d9f5cc284b · outbound

This paper cites surrogate advantage trick.

Central Path Proximal Policy Optimization surrogate advantage trick

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.352342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:09.566879Z digest=sha256:9c18c28266d792c539c34d56c13809f73a20095959b6e9e2f6803cd914ed26ff

Observation da807287-27f3-458c-99d7-ddad26648efb · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Central Path Proximal Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.141595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.141595Z digest=sha256:4e6081651540f7f66a7edcef063b55077f9b553ef222f52b8b12ad32af21538b

Observation d1d34400-fb63-46a9-b893-1e669deacd40 · outbound

This paper cites Dimitri P Bertsekas.

Central Path Proximal Policy Optimization Dimitri P Bertsekas

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.664488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:08.578032Z digest=sha256:35ec6bfda7101dd99d858a07ef2ba48b1ecba00495f81b1497f8575c2debd8a9

Observation 3f78ce86-e4d7-4720-8ad6-c368ea7bfc38 · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

Central Path Proximal Policy Optimization Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.633497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.633497Z digest=sha256:b9788d802683f4c2d25e587c732915a016d71620918299e7e8c1f2b0d9472b95

Observation 2f5a27ad-336f-436e-9782-1670098eb2d9 · outbound

This paper cites Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans.

Central Path Proximal Policy Optimization Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.702350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.702350Z digest=sha256:4a13f9d46e451b132b090399c3f477ff73b404738a3b252eeb400227b0667b5a

Observation 6ba153c4-cc06-4dc8-bf5c-d9d8f999b133 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Central Path Proximal Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.990567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.990567Z digest=sha256:70c57376470f78a76f6167bc787a4f03817a593d59555195c7ae54f396b93844

Observation 4b21ca17-83c7-407c-bef8-f3c3fc4823e1 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Central Path Proximal Policy Optimization A unified view of entropy-regularized Markov decision processes

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.844809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.844809Z digest=sha256:6aeb5f01d4bd8994ee4b0b121c8393718070a604c997f25d356242ad2daa70e1

Observation 9eb6c714-6dfd-4a21-b04b-ee6f1af7c67d · outbound

This paper cites On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization.

Central Path Proximal Policy Optimization On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.017454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:07:09.219280Z digest=sha256:95a82d0f1cec75ea5ca137e17f97822a3edd0c9bb10dc656a5d41255c93def1b

Observation fc27e03e-c0c5-4d02-a012-6a3d05bf8902 · outbound

This paper cites Embedding Safety into RL: A New Take on Trust Region Methods.

Central Path Proximal Policy Optimization Embedding Safety into RL: A New Take on Trust Region Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.760132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.760132Z digest=sha256:549bdcbe6f3a89180d79c087ca276cdafb09b480f880c502dd69a3e93a2530dc

Pith citing papers

Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · inbound

The Geometry of Nonlinear Reinforcement Learning cites this paper.

The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.630698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.630698Z digest=sha256:8abc165b1052c09973d4d61e97a2c00073e6fd22d6c068395fa9ab25a80fa15c

Observation d095a95e-6a3c-48b7-817c-c5acf95f3998 · inbound

Bounded Ratio Reinforcement Learning cites this paper.

Bounded Ratio Reinforcement Learning Central Path Proximal Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:19.079386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:50:11.020901Z digest=sha256:1e79a0699aa15e7f82c074ea8a883e766b89a13862069c9c0bb66ac685a240c3