Pith. sign in

Paper Citation Record · LEDGER

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.10204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10204 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:18.685092Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10b10e98-50eb-49bb-9ce7-81c21fff59e3 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.611611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.611611Z digest=sha256:9679a5e074a1d902ae1e24a2d5499868532f1e0dc0f0b9e3e9bcd89d897de7c4

Observation 2e38e038-6f10-444e-bcd8-fae6c37ef433 · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.638317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.638317Z digest=sha256:e3a834bab8249e61798b0817fe72aa5bb3a276c694049f6e9138201e2200c53f

Observation a9926449-e9f5-4944-ad6c-7542f4c73c04 · outbound

This paper cites Reinforcement learning in robotics: A survey,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reinforcement learning in robotics: A survey,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.686177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.686177Z digest=sha256:57a6f48cbda78852f3af53e067b71a5fa83bc30f295871ef8ef3029c1ffdc04a

Observation a7f3e378-c7d3-40bf-b3b8-1f3914037b05 · outbound

This paper cites Altman,Constrained Markov decision processes.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Altman,Constrained Markov decision processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.734749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.734749Z digest=sha256:ce31d0e8515fddd292acc95e3cb0b36653dc812c3cbeed6ebf1fe4d5763780b8

Observation 74ea29d0-8faa-42b9-b91e-27589e0afd6a · outbound

This paper cites Responsive safety in rein- forcement learning by pid lagrangian methods,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Responsive safety in rein- forcement learning by pid lagrangian methods,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.386367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:17.784810Z digest=sha256:08884ce26d92498e166bb7c2ac6cb94b8e7fdae1290e6f53af3c2202c6cb42af

Observation 5cff693b-2d6a-4cf2-835b-7489c820a726 · outbound

This paper cites Safe policies for reinforcement learning via primal-dual methods,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe policies for reinforcement learning via primal-dual methods,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.303537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:17.844751Z digest=sha256:7c54097c934bff7bc77efff8cdf7ae6bde22fc0ee61ffea3053d5fa7d4fa7b9a

Observation 12d1574a-c11f-4233-a768-21f351463218 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.285379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:17.904239Z digest=sha256:cd76bee2fe770214b05855a4dc35c4259b7242508ac677a088f2723d234679d4

Observation 0a6ec3c6-6c11-4b31-a68f-a4468e6dab8c · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDPs,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient model-free algorithms for non-stationary CMDPs,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.254746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:17.954758Z digest=sha256:3e746f795a7c98a68211edf77088270b5ce4fbb3344e8541eca5fa84eaaa06e8

Observation b0e41404-8994-4b66-b155-9cf607fe4701 · outbound

This paper cites Safe and efficient: A primal- dual method for offline convex cmdps under partial data coverage,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe and efficient: A primal- dual method for offline convex cmdps under partial data coverage,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.154076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.002316Z digest=sha256:9445bece006d919627dc4f8d9e5255d870f945b65853fd077269a86026307387

Observation 2295f152-e0d3-4459-a2de-2ca8ba851e90 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient safe exploration via primal-dual policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.091842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.044752Z digest=sha256:2d5de7fa3a4009ad2b58ae98f28d0b723bae1603209f62fbf88f6005772860c0

Observation 8a3387d2-5153-44f6-afff-4116291a0503 · outbound

This paper cites An optimistic algorithm for online cmdps with anytime adversarial constraints,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning An optimistic algorithm for online cmdps with anytime adversarial constraints,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.016276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.074748Z digest=sha256:d5e78b6b3f0821832e07722afd0ff41e48cd62b08283685bf0f96999ec9b1044

Observation 5c81a8c8-5790-4374-b34f-bf7ee6d0a1f2 · outbound

This paper cites Constrained policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained policy optimization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.987125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.114750Z digest=sha256:8c53fd619c41053c140ca38d4878cea9a8bcf5c8be73fd32049f68ba8a1e0cb3

Observation d61b0c7e-0248-452c-a3e2-824191f92c3f · outbound

This paper cites Projection- based constrained policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Projection- based constrained policy optimization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.935449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.154736Z digest=sha256:6c5a237315be1bdb316c308c06076620b348057b54a1fd8aa1769e0ad41811b5

Observation a96b4ffa-d668-493f-8705-5a53d89cbfe2 · outbound

This paper cites First order constrained optimization in policy space,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning First order constrained optimization in policy space,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.874747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.174739Z digest=sha256:e171c6b8ff3c07bc8a0989f3f6b886e178e2d167eff9c8e73b3c45b576d81dc9

Observation 1773ba1a-a689-4e7e-9ba5-1698dd5d63df · outbound

This paper cites Constrained update projection approach to safe policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained update projection approach to safe policy optimization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.787414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.218045Z digest=sha256:bdb7ad8de241d606e626d9fb3417b6860ca7fb69b85eda1df5867b253bb92a12

Observation 5fbecf50-96b8-4fd1-956a-6560ee042522 · outbound

This paper cites Crpo: A new approach for safe reinforcement learning with convergence guarantee,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Crpo: A new approach for safe reinforcement learning with convergence guarantee,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.744760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.252129Z digest=sha256:0d61c39c905a37dcaa344d8bae5637802b6d94a228ab29c120359a75504cb008

Observation d00a5707-5713-4ac9-a1b3-849d61ddb0f8 · outbound

This paper cites Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.688124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.295917Z digest=sha256:01b0a1d17efbfda61547495d6084db9b5bcab3750cb1ad5b0618de15c517f8e0

Observation 4d803e81-1e17-4e34-9aad-c8e2dca5c8bd · outbound

This paper cites Enhancing efficiency of safe reinforcement learning via sample manipulation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Enhancing efficiency of safe reinforcement learning via sample manipulation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.618194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.313396Z digest=sha256:285b332f9c7950eea7943f93ac0eb42eefd156f2f94abb715a54840476c39c7b

Observation c1ff5d35-686b-4895-96f6-257a6b804b5a · outbound

This paper cites Gradient shaping for multi-constraint safe reinforcement learning,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Gradient shaping for multi-constraint safe reinforcement learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.519688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.333619Z digest=sha256:165cd42e7652b9bbe5cd938f4ce7c8e075e08d065fa7a9fac9a4c16f592f236e

Observation 25f95649-afec-4334-8810-b788610e90f5 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.373663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.373663Z digest=sha256:3ff46d4942661af60c638fb930bc3bdb9d7c0011c9e70ca940c64d3b98fb8b91

Observation d913a37e-862d-42ad-9ecb-da71dcc6bc99 · outbound

This paper cites Reward Constrained Policy Optimization.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reward Constrained Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.414100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.414100Z digest=sha256:88e0f001836216de3f657acf1955151294767b3034b3c17307834008d8b633fa

Observation 70f41fff-2e27-4050-a97e-5e3821b9ef89 · outbound

This paper cites Off-Policy Primal-Dual Safe Reinforcement Learning.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Off-Policy Primal-Dual Safe Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.465113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.465113Z digest=sha256:f2162059f1d8c364ca8cec80fb8a4664ee842af3d3151adecbe9910f31b69686

Observation d7f6727f-04c3-4f0f-965a-6c3a705457ba · outbound

This paper cites Constrained Proximal Policy Optimization.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained Proximal Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.507270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.507270Z digest=sha256:bcfb79f87faa89d05911b0fb2a1d5d382febb1eb6a98d34cc4148aea7035b5dd

Observation f1a80751-9b09-418b-b1ec-8a866bb4d9c9 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.419344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.555032Z digest=sha256:b7c0559794f31dce493233b44cd35a13a536fb0e5ad8d6c7bc999b10901da449

Observation c95c6dd2-3f9a-49a4-b2cd-ff258b8a4565 · outbound

This paper cites Proximal policy optimization algorithms,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Proximal policy optimization algorithms,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.605434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.605434Z digest=sha256:97887eea55a6748b2e2a10a4ec7bf09f2d813befbea2a00acd9edc58b4319d06

Observation a2f22536-bd7b-408d-ba36-ec556e44e324 · outbound

This paper cites an unresolved cited work.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.654752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.654752Z digest=sha256:df90ec99c4766e18ce8c8af89674401db051526514ba73564299a6912e71b8c7

Observation fd535ee1-ffba-435b-a9f9-64d464da01ac · outbound

This paper cites Omnisafe: An infrastructure for accelerating safe reinforcement learning research,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Omnisafe: An infrastructure for accelerating safe reinforcement learning research,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.204882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T04:18:18.685092Z digest=sha256:8b69444bf64addf7ccb855d8d096aa52795195a7e901fe2cf9ee56adc6c9bcac

Pith citing papers

No inbound Pith citation observations are available.