Pith. sign in

Paper Citation Record · LEDGER

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.10204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10204 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:18.685092Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10b10e98-50eb-49bb-9ce7-81c21fff59e3 · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.611611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.611611Z digest=sha256:9679a5e074a1d902ae1e24a2d5499868532f1e0dc0f0b9e3e9bcd89d897de7c4

Observation 2e38e038-6f10-444e-bcd8-fae6c37ef433 · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.638317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.638317Z digest=sha256:e3a834bab8249e61798b0817fe72aa5bb3a276c694049f6e9138201e2200c53f

Observation a9926449-e9f5-4944-ad6c-7542f4c73c04 · outbound

This paper cites Reinforcement learning in robotics: A survey,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reinforcement learning in robotics: A survey,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.686177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.686177Z digest=sha256:57a6f48cbda78852f3af53e067b71a5fa83bc30f295871ef8ef3029c1ffdc04a

Observation a7f3e378-c7d3-40bf-b3b8-1f3914037b05 · outbound

This paper cites Altman,Constrained Markov decision processes.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Altman,Constrained Markov decision processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:17.734749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:17.734749Z digest=sha256:ce31d0e8515fddd292acc95e3cb0b36653dc812c3cbeed6ebf1fe4d5763780b8

Observation 74ea29d0-8faa-42b9-b91e-27589e0afd6a · outbound

This paper cites Responsive safety in rein- forcement learning by pid lagrangian methods,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Responsive safety in rein- forcement learning by pid lagrangian methods,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.386367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:17.784810Z digest=sha256:31917f41e7da9af320b95287a59807655db406c9e066753e406e5e225f9b6d24

Observation 5cff693b-2d6a-4cf2-835b-7489c820a726 · outbound

This paper cites Safe policies for reinforcement learning via primal-dual methods,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe policies for reinforcement learning via primal-dual methods,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.303537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:17.844751Z digest=sha256:047c486b1f8fcd5712f30e3ae559e6d634a1ec8d5480fe7555fa9df3d3dd47bd

Observation 12d1574a-c11f-4233-a768-21f351463218 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.285379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:17.904239Z digest=sha256:5c6f22a0b82bce2f127277853412652fa5466574eaf5f22af87eb66018b6c708

Observation 0a6ec3c6-6c11-4b31-a68f-a4468e6dab8c · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDPs,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient model-free algorithms for non-stationary CMDPs,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.254746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:17.954758Z digest=sha256:08d46b155bb48c95acae0a9618cef883adc364f54c775e6eee0f6043f25b92cd

Observation b0e41404-8994-4b66-b155-9cf607fe4701 · outbound

This paper cites Safe and efficient: A primal- dual method for offline convex cmdps under partial data coverage,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Safe and efficient: A primal- dual method for offline convex cmdps under partial data coverage,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.154076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.002316Z digest=sha256:e255f3cfbc72a9ab44666b215e7d6c5c12c1675f6697cbfcd91242446caac9fd

Observation 2295f152-e0d3-4459-a2de-2ca8ba851e90 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Provably efficient safe exploration via primal-dual policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.091842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.044752Z digest=sha256:c71435e2337f435753a489a1db6f7061310e4d1cc9bb76bffeb24b61f1e7d2fc

Observation 8a3387d2-5153-44f6-afff-4116291a0503 · outbound

This paper cites An optimistic algorithm for online cmdps with anytime adversarial constraints,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning An optimistic algorithm for online cmdps with anytime adversarial constraints,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:20.016276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.074748Z digest=sha256:f70d61718adc7975fb634c0414d0851cc49ede3bfd6f508adf674a4596e4bb6e

Observation 5c81a8c8-5790-4374-b34f-bf7ee6d0a1f2 · outbound

This paper cites Constrained policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained policy optimization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.987125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.114750Z digest=sha256:db889093fe27dae07cfa3be8a410a53c1467e3485938b47b3e4789c1aaff8d2e

Observation d61b0c7e-0248-452c-a3e2-824191f92c3f · outbound

This paper cites Projection- based constrained policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Projection- based constrained policy optimization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.935449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.154736Z digest=sha256:51e8dd81e28df40eb2808bf7cfe1fd7845ef8363d9364defe03456f6bc1602e0

Observation a96b4ffa-d668-493f-8705-5a53d89cbfe2 · outbound

This paper cites First order constrained optimization in policy space,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning First order constrained optimization in policy space,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.874747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.174739Z digest=sha256:cbbb924b09a6b12caea455d639e6751e76886cc115cef1e11e8305fe45f91728

Observation 1773ba1a-a689-4e7e-9ba5-1698dd5d63df · outbound

This paper cites Constrained update projection approach to safe policy optimization,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained update projection approach to safe policy optimization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.787414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.218045Z digest=sha256:bab939f13d4531b6002ea24864149292a82571bf2469e077956dfe7473f78ab3

Observation 5fbecf50-96b8-4fd1-956a-6560ee042522 · outbound

This paper cites Crpo: A new approach for safe reinforcement learning with convergence guarantee,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Crpo: A new approach for safe reinforcement learning with convergence guarantee,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.744760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.252129Z digest=sha256:d5c61b48cba60ac4775463166cf88008fdcaa86bd874b4f39e537d25e921e0e1

Observation d00a5707-5713-4ac9-a1b3-849d61ddb0f8 · outbound

This paper cites Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.688124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.295917Z digest=sha256:46479d0d2fe364a87c9946938aa8d272897248417da7afc390c11d550dea639f

Observation 4d803e81-1e17-4e34-9aad-c8e2dca5c8bd · outbound

This paper cites Enhancing efficiency of safe reinforcement learning via sample manipulation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Enhancing efficiency of safe reinforcement learning via sample manipulation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.618194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.313396Z digest=sha256:5a5453163ec8c1f215d1fa87cbf439f385ad94de756a889e36fce5bacd463e19

Observation c1ff5d35-686b-4895-96f6-257a6b804b5a · outbound

This paper cites Gradient shaping for multi-constraint safe reinforcement learning,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Gradient shaping for multi-constraint safe reinforcement learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.519688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.333619Z digest=sha256:7a1e566456b431b5807da0fe775edddc71f83bcd9bd5b91de9ed48f0a8ce5523

Observation 25f95649-afec-4334-8810-b788610e90f5 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.373663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.373663Z digest=sha256:3ff46d4942661af60c638fb930bc3bdb9d7c0011c9e70ca940c64d3b98fb8b91

Observation d913a37e-862d-42ad-9ecb-da71dcc6bc99 · outbound

This paper cites Reward Constrained Policy Optimization.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Reward Constrained Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.414100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.414100Z digest=sha256:88e0f001836216de3f657acf1955151294767b3034b3c17307834008d8b633fa

Observation 70f41fff-2e27-4050-a97e-5e3821b9ef89 · outbound

This paper cites Off-Policy Primal-Dual Safe Reinforcement Learning.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Off-Policy Primal-Dual Safe Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.465113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.465113Z digest=sha256:4f67e27fd40b431616fca88de65214691e2ce94a9e52043f866df106972cc73a

Observation d7f6727f-04c3-4f0f-965a-6c3a705457ba · outbound

This paper cites Constrained Proximal Policy Optimization.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Constrained Proximal Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.507270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.507270Z digest=sha256:5f90b26f8d367d2fa8e0d68b8b2953be9472b6e0d2b67c0e8f9695e418552e81

Observation f1a80751-9b09-418b-b1ec-8a866bb4d9c9 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.419344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.555032Z digest=sha256:3e83cef5f6bc5e5050062cd899af060f8f1dca6373f1e343d065d0ee23841385

Observation c95c6dd2-3f9a-49a4-b2cd-ff258b8a4565 · outbound

This paper cites Proximal policy optimization algorithms,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Proximal policy optimization algorithms,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.605434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.605434Z digest=sha256:97887eea55a6748b2e2a10a4ec7bf09f2d813befbea2a00acd9edc58b4319d06

Observation a2f22536-bd7b-408d-ba36-ec556e44e324 · outbound

This paper cites an unresolved cited work.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.654752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.654752Z digest=sha256:df90ec99c4766e18ce8c8af89674401db051526514ba73564299a6912e71b8c7

Observation fd535ee1-ffba-435b-a9f9-64d464da01ac · outbound

This paper cites Omnisafe: An infrastructure for accelerating safe reinforcement learning research,.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Omnisafe: An infrastructure for accelerating safe reinforcement learning research,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:19.204882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:18:18.685092Z digest=sha256:91ea7fda4794804ba9946d7dcdd844b3d0a9acbdaf6731c72d355e4c72f1793b

Pith citing papers

No inbound Pith citation observations are available.