Pith. sign in

Paper Citation Record · LEDGER

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2509.09208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09208 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:35:33.330106Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95a53441-1ef7-442e-bc2a-cb941902f988 · outbound

This paper cites Constrained policy optimiza- tion.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Constrained policy optimiza- tion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.235570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.235570Z digest=sha256:46faafec55c0d894f1ac77134f090dbeb2a2a581e7e4460b79d756a03bc423e6

Observation 52a8f122-addf-4e2d-b04a-435b0103722e · outbound

This paper cites We also conducted experi- ments using the MetaDrive simulator [Liet al., 2022 ].

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning We also conducted experi- ments using the MetaDrive simulator [Liet al., 2022 ]

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.330106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.330106Z digest=sha256:dc62749be9950feb576bd1f1f09ac5a1d78b46deaf8985851634f2519fd8e1b1

Observation 92f6143e-f2d5-46ba-81f8-651d4a92e068 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.Journal of Machine Learning Research, 18(167):1–51,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria.Journal of Machine Learning Research, 18(167):1–51,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.245700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.245700Z digest=sha256:ba7fa4f49c1e43431ec90d12884576bed26453ae51c3a908ae16d3b88afc3b53

Observation eb978cd1-26c8-4385-b799-bdafa0f612b5 · outbound

This paper cites A general safety framework for learning-based control in uncertain robotic systems.IEEE Transactions on Automatic Control, 64(7):2737–2752,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning A general safety framework for learning-based control in uncertain robotic systems.IEEE Transactions on Automatic Control, 64(7):2737–2752,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.259379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.259379Z digest=sha256:b9f43ddbd0e1c16f4dc52e4e61d4d240d2e42302b907d95a4980939de0b29dd2

Observation 66e6ae51-5f20-428c-9742-dab8e260c8b0 · outbound

This paper cites Exterior penalty pol- icy optimization with penalty metric network under con- straints.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Exterior penalty pol- icy optimization with penalty metric network under con- straints

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.261426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.261426Z digest=sha256:5198b3c183d7141dc8d9e8f43679daea8993b91f6431a14ac30637541ff3afd1

Observation 68734fb2-01db-4e23-bae1-dec886263353 · outbound

This paper cites Omnisafe: An infrastructure for accelerating safe reinforcement learn- ing research.Journal of Machine Learning Research, 25(285):1–6,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Omnisafe: An infrastructure for accelerating safe reinforcement learn- ing research.Journal of Machine Learning Research, 25(285):1–6,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.268094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.268094Z digest=sha256:217c579ff7cf400eb8dcf5946a575bf3ac43ddbad6b68664000ba9ce13d84a6c

Observation 328d2aea-eae0-4ec4-8c95-871cb8c32ff6 · outbound

This paper cites Doubly ro- bust off-policy value evaluation for reinforcement learn- ing.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Doubly ro- bust off-policy value evaluation for reinforcement learn- ing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.270031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.270031Z digest=sha256:10572c5a87d2ae6f473cbd8257c3608c258cd165639755e6e34616a0a2954138

Observation 252ec687-fbce-4651-8963-c9ce3940e728 · outbound

This paper cites End-to-end training of deep visuomotor policies.Journal of Machine Learning Re- search, 17(39):1–40,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning End-to-end training of deep visuomotor policies.Journal of Machine Learning Re- search, 17(39):1–40,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.276785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.276785Z digest=sha256:aee18a9de6927d80c78eb39b10bc66f7447c1f847676eb80e549d32449da74af

Observation 93993193-d870-47b0-916b-140b5c9eeec9 · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable re- inforcement learning.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3461–3475,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Metadrive: Composing diverse driving scenarios for generalizable re- inforcement learning.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3461–3475,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.278849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.278849Z digest=sha256:9b445200788515d65de6ec540fab766fe36db43f17e7163cd1e014e32978f446

Observation 92526f50-13b8-4949-bc85-2053192b2e3a · outbound

This paper cites Ipo: Interior-point policy optimization under constraints.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Ipo: Interior-point policy optimization under constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.280812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.280812Z digest=sha256:c827e25e32c0e6370ec7184cc23ef342f4c3bee93104c0dc0d3f242845d0bec7

Observation 03155a60-652d-48d2-a8f3-1b2858e230cc · outbound

This paper cites The law and ethics of high-frequency trading.Minn.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning The law and ethics of high-frequency trading.Minn

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.283481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.283481Z digest=sha256:9e20153d48f17d634746ece2d4d2423ab9d0711f68216fdc562eb6103dd6cd57

Observation 4316231b-44a0-4c3d-a49f-dcffc2de81c3 · outbound

This paper cites Chance-constrained dynamic pro- gramming with application to risk-aware robotic space ex- ploration.Autonomous Robots, 39:555–571,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Chance-constrained dynamic pro- gramming with application to risk-aware robotic space ex- ploration.Autonomous Robots, 39:555–571,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.286703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.286703Z digest=sha256:6f03fa00e9e43a2884fcc8d908b4d84d0b0cd4e7bacd5c28c62b92d2ff5b1c1e

Observation 1598ab19-b2c5-43fb-82f7-c67e9faca4b6 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.289482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.289482Z digest=sha256:94f63d9fccd7ce9691f6e2eedd6add0848fb194f8bb92e488c8d12517d75c6f9

Observation 59ee5433-572a-4d72-bcd9-7fc6a633777e · outbound

This paper cites Trust Region Policy Optimization.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Trust Region Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.294233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.294233Z digest=sha256:dfe6f2abd3c626dedec00256c8e0a45ab467669a54540c88a6da0e613fec41c1

Observation d20978f6-ef78-4a88-9f84-250785906017 · outbound

This paper cites Responsive safety in reinforcement learn- ing by pid lagrangian methods.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Responsive safety in reinforcement learn- ing by pid lagrangian methods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.296483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.296483Z digest=sha256:0747da4cfde36aa228fb2111dbc0d99914cb1631e396612a23413e26e44fa0eb

Observation 1e281c07-29d1-42fe-9fe2-7bc6451d9406 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.299036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.299036Z digest=sha256:ad09696c9d204e284106398d2cd9f515791c34b50510e8c863d2093e17a0e0b0

Observation 89af7791-721c-43ec-81d8-3ff9c9f8001b · outbound

This paper cites MIT press,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning MIT press,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.301563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.301563Z digest=sha256:ea85acdc1106cd424a0c43e53b662e97ffe5f00c5670f0e0295906035c960119

Observation 0af87724-2aea-4df1-b2ed-7a98700a2a8c · outbound

This paper cites Reward constrained policy optimization.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Reward constrained policy optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.304615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.304615Z digest=sha256:d80b6f51813f568b019c51a8fd329c4024133d0f87f83334ebbf066b0e047d1c

Observation 3da88d1b-5293-4cf6-a9c3-8a68470baecf · outbound

This paper cites Projection- based constrained policy optimization.International Conference on Learning Representations,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Projection- based constrained policy optimization.International Conference on Learning Representations,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.306999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.306999Z digest=sha256:ef4aa5ce2db2e1e7d72ab9a045e7506aaf1adb4544b0861a72baa6a0197e08d2

Observation 2b6ffabd-050d-47bf-a739-c4ea6c7affe7 · outbound

This paper cites Constrained update projection approach to safe policy optimization.Advances in Neural Information Pro- cessing Systems, 35:9111–9124,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Constrained update projection approach to safe policy optimization.Advances in Neural Information Pro- cessing Systems, 35:9111–9124,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.309374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.309374Z digest=sha256:e7135c7869fa5448d7b5b7f049affcd7d5908b9b3eb51e82c6cd4c97b8989a48

Observation b3188aa4-c4f0-41ca-a5ba-8d013d383bcf · outbound

This paper cites Convergent policy optimization for safe reinforcement learning.Advances in Neural Information Processing Systems, 32,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Convergent policy optimization for safe reinforcement learning.Advances in Neural Information Processing Systems, 32,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.311273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.311273Z digest=sha256:261d742f77ce2d2e52dc94fdc98f3af282212db0fe67135f7dbc84bfab28e434

Observation 700e0552-a5f6-4792-93ef-7b4799ff5700 · outbound

This paper cites Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.313763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.313763Z digest=sha256:53e23b8c99f7d94184d1378211f6e8596ca0e143464b4b55a7034e75fe30a6bd

Observation 8c4942c9-1fbf-44c3-8412-807359aa70f2 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.Advances in Neural Information Processing Sys- tems, 35:24611–24624,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning The surprising effectiveness of ppo in cooperative multi-agent games.Advances in Neural Information Processing Sys- tems, 35:24611–24624,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.316231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.316231Z digest=sha256:6d734a33facfc24e4e361b926a5e1d790d6ff76c7820cf5c7d68d054976fd063

Observation aa8c2496-165a-4616-8029-4d8d2ffdb9ff · outbound

This paper cites First order constrained optimization in policy space.Advances in Neural Information Processing Sys- tems, 33:15338–15349,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning First order constrained optimization in policy space.Advances in Neural Information Processing Sys- tems, 33:15338–15349,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.318854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.318854Z digest=sha256:6003450d065bd17726181de68408cfe4045769ddfbc771f93575acd06ec2bfdd

Observation bb8119d5-a7cb-4654-81f8-785fa9f46316 · outbound

This paper cites Penalized proximal policy optimization for safe rein- forcement learning.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Penalized proximal policy optimization for safe rein- forcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.320745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.320745Z digest=sha256:a3bb9c07cb428711c40f6ef05a62eaf28b63c9bb6d58c3b719f51e56bc9fa4be

Observation 90ad360b-aae8-4b5f-9216-ebe30b0d1246 · outbound

This paper cites Evaluating model-free reinforcement learning toward safety-critical tasks.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.323026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.323026Z digest=sha256:cdaba0ef2a0a258b8497f67b1ae1feace1943e1c05994163d77cb194c63f039f

Observation 5d9a9963-f392-473d-8c3f-0aba7a6b5bd4 · outbound

This paper cites an unresolved cited work.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.325727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.325727Z digest=sha256:a3401e31ab9504c201b69c076dc104a0bcde4113f813ff90959d769465ed785c

Observation a380bf07-c3c3-4ec8-ae61-ac568aa8ef55 · outbound

This paper cites an unresolved cited work.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.328059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.328059Z digest=sha256:0ab1ecccd31e3a2a884c517c8fd98e026a4fce133e6ce4b3591edb5b5621cbd5

Observation b9ebec16-e3f4-49ef-88fc-7a321e6128bc · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.IEEE Transactions on In- telligent Transportation Systems, 23(6):4909–4926,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Deep reinforcement learning for autonomous driving: A survey.IEEE Transactions on In- telligent Transportation Systems, 23(6):4909–4926,

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.274485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.274485Z digest=sha256:66aaacb90c1b4ef979a7f7ea11fb86b1131323a3d1ca000d1529bd55a1601bf5

Observation 2fae1680-ee90-429f-8558-5b3e8b73d0a9 · outbound

This paper cites Augmented proximal policy op- timization for safe reinforcement learning.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Augmented proximal policy op- timization for safe reinforcement learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.250837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.250837Z digest=sha256:f7b9e5f3036916aadffaab75f14265ce670fb516bcc166baccac33116c07a6bc

Observation a952f2ee-ef59-40ac-a35e-c9e192aad329 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Approximately optimal approximate reinforcement learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.272178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.272178Z digest=sha256:90976283791564a715220c69be3684aa71a294c624b1b4a8f16524e32bf6ee16

Observation f0d0627e-ccd6-4b22-a505-a29acc84d432 · outbound

This paper cites Routledge,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Routledge,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.239265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.239265Z digest=sha256:f41ed29069e0c8e93f6f28543e6f5f7f9db9d29a61dd495060c22c3a92da7162

Observation b407b949-5a32-4bfd-bf24-2a9058e53444 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.248247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.248247Z digest=sha256:11f4528532700546d54ec6177ac5b3a6716b14e0328c943882c59745946e4b41

Observation 810483af-9f18-4f90-9d26-a9093c06bbea · outbound

This paper cites Proximal Policy Optimization Algorithms.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.291635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.291635Z digest=sha256:f32442152e14654b7752878a8dc9e9a756ebb43f500c8220fc12efc0f856c327

Observation d54d7a10-f8a2-4646-8b43-a66f2cf5d361 · outbound

This paper cites Trustworthy artificial intelligence requirements in the autonomous driving domain.Computer, 56(2):29–39,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Trustworthy artificial intelligence requirements in the autonomous driving domain.Computer, 56(2):29–39,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.256528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.256528Z digest=sha256:6d23dcc6395f8e85813f8605675f530fec86d7bd660c3433b15fe75ff30f2dc0

Observation 028dc12d-cf81-4fe9-9fea-c2e75bbae4e8 · outbound

This paper cites Continuously Differentiable Exponential Linear Units.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Continuously Differentiable Exponential Linear Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.242544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.242544Z digest=sha256:3aad3ed1d531e0fea6468e74d912228bec05b6e7d906c3d01ed012e3f8c4bb2e

Observation 67742a80-741a-421e-9d6b-3ffee382841c · outbound

This paper cites Safety gymna- sium: A unified safe reinforcement learning benchmark.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.265931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.265931Z digest=sha256:686479a6fd5e69b6b45f05236ddcb8293ca64dc09a222eb0fece5cef05931c83

Observation 17694841-25c3-4539-942b-48e915c3b10d · outbound

This paper cites Natural policy gradient primal-dual method for constrained markov decision pro- cesses.Advances in Neural Information Processing Sys- tems, 33:8378–8390,.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Natural policy gradient primal-dual method for constrained markov decision pro- cesses.Advances in Neural Information Processing Sys- tems, 33:8378–8390,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.254104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.254104Z digest=sha256:58fb4bb05d4220555caccd29c23615ec4903945146ef9404c5adc997b2c6e1e3

Observation 440be212-e38c-4ca0-a9a5-d1b7712cbf71 · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.263507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.263507Z digest=sha256:f96e7f89cc97325acf497f1b902510a89f7b1a551c8d265aa43465f0b065d104

Pith citing papers

No inbound Pith citation observations are available.