Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:18.685092Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.10204.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:18.685092Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 10b10e98-50eb-49bb-9ce7-81c21fff59e3 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e38e038-6f10-444e-bcd8-fae6c37ef433 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9926449-e9f5-4944-ad6c-7542f4c73c04 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reinforcement learning in robotics: A survey,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f3e378-c7d3-40bf-b3b8-1f3914037b05 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Altman,Constrained Markov decision processes
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ea29d0-8faa-42b9-b91e-27589e0afd6a · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Responsive safety in rein- forcement learning by pid lagrangian methods,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5cff693b-2d6a-4cf2-835b-7489c820a726 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe policies for reinforcement learning via primal-dual methods,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 12d1574a-c11f-4233-a768-21f351463218 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a6ec3c6-6c11-4b31-a68f-a4468e6dab8c · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient model-free algorithms for non-stationary CMDPs,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b0e41404-8994-4b66-b155-9cf607fe4701 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe and efficient: A primal- dual method for offline convex cmdps under partial data coverage,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2295f152-e0d3-4459-a2de-2ca8ba851e90 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient safe exploration via primal-dual policy optimization,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a3387d2-5153-44f6-afff-4116291a0503 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning An optimistic algorithm for online cmdps with anytime adversarial constraints,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c81a8c8-5790-4374-b34f-bf7ee6d0a1f2 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained policy optimization,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d61b0c7e-0248-452c-a3e2-824191f92c3f · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Projection- based constrained policy optimization,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a96b4ffa-d668-493f-8705-5a53d89cbfe2 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning First order constrained optimization in policy space,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1773ba1a-a689-4e7e-9ba5-1698dd5d63df · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained update projection approach to safe policy optimization,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5fbecf50-96b8-4fd1-956a-6560ee042522 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Crpo: A new approach for safe reinforcement learning with convergence guarantee,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d00a5707-5713-4ac9-a1b3-849d61ddb0f8 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4d803e81-1e17-4e34-9aad-c8e2dca5c8bd · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Enhancing efficiency of safe reinforcement learning via sample manipulation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c1ff5d35-686b-4895-96f6-257a6b804b5a · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Gradient shaping for multi-constraint safe reinforcement learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 25f95649-afec-4334-8810-b788610e90f5 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d913a37e-862d-42ad-9ecb-da71dcc6bc99 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reward Constrained Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f41fff-2e27-4050-a97e-5e3821b9ef89 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Off-Policy Primal-Dual Safe Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f6727f-04c3-4f0f-965a-6c3a705457ba · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained Proximal Policy Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a80751-9b09-418b-b1ec-8a866bb4d9c9 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c95c6dd2-3f9a-49a4-b2cd-ff258b8a4565 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Proximal policy optimization algorithms,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f22536-bd7b-408d-ba36-ec556e44e324 · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd535ee1-ffba-435b-a9f9-64d464da01ac · outbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Omnisafe: An infrastructure for accelerating safe reinforcement learning research,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.