Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:35:33.330106Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2509.09208.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:35:33.330106Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95a53441-1ef7-442e-bc2a-cb941902f988 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Constrained policy optimiza- tion
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a8f122-addf-4e2d-b04a-435b0103722e · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning We also conducted experi- ments using the MetaDrive simulator [Liet al., 2022 ]
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f6143e-f2d5-46ba-81f8-651d4a92e068 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria.Journal of Machine Learning Research, 18(167):1–51,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb978cd1-26c8-4385-b799-bdafa0f612b5 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning A general safety framework for learning-based control in uncertain robotic systems.IEEE Transactions on Automatic Control, 64(7):2737–2752,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e6ae51-5f20-428c-9742-dab8e260c8b0 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Exterior penalty pol- icy optimization with penalty metric network under con- straints
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68734fb2-01db-4e23-bae1-dec886263353 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Omnisafe: An infrastructure for accelerating safe reinforcement learn- ing research.Journal of Machine Learning Research, 25(285):1–6,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328d2aea-eae0-4ec4-8c95-871cb8c32ff6 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Doubly ro- bust off-policy value evaluation for reinforcement learn- ing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252ec687-fbce-4651-8963-c9ce3940e728 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning End-to-end training of deep visuomotor policies.Journal of Machine Learning Re- search, 17(39):1–40,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93993193-d870-47b0-916b-140b5c9eeec9 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Metadrive: Composing diverse driving scenarios for generalizable re- inforcement learning.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3461–3475,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92526f50-13b8-4949-bc85-2053192b2e3a · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Ipo: Interior-point policy optimization under constraints
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03155a60-652d-48d2-a8f3-1b2858e230cc · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning The law and ethics of high-frequency trading.Minn
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4316231b-44a0-4c3d-a49f-dcffc2de81c3 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Chance-constrained dynamic pro- gramming with application to risk-aware robotic space ex- ploration.Autonomous Robots, 39:555–571,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1598ab19-b2c5-43fb-82f7-c67e9faca4b6 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ee5433-572a-4d72-bcd9-7fc6a633777e · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Trust Region Policy Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20978f6-ef78-4a88-9f84-250785906017 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Responsive safety in reinforcement learn- ing by pid lagrangian methods
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e281c07-29d1-42fe-9fe2-7bc6451d9406 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89af7791-721c-43ec-81d8-3ff9c9f8001b · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning MIT press,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af87724-2aea-4df1-b2ed-7a98700a2a8c · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Reward constrained policy optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da88d1b-5293-4cf6-a9c3-8a68470baecf · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Projection- based constrained policy optimization.International Conference on Learning Representations,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6ffabd-050d-47bf-a739-c4ea6c7affe7 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Constrained update projection approach to safe policy optimization.Advances in Neural Information Pro- cessing Systems, 35:9111–9124,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3188aa4-c4f0-41ca-a5ba-8d013d383bcf · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Convergent policy optimization for safe reinforcement learning.Advances in Neural Information Processing Systems, 32,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700e0552-a5f6-4792-93ef-7b4799ff5700 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c4942c9-1fbf-44c3-8412-807359aa70f2 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning The surprising effectiveness of ppo in cooperative multi-agent games.Advances in Neural Information Processing Sys- tems, 35:24611–24624,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8c2496-165a-4616-8029-4d8d2ffdb9ff · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning First order constrained optimization in policy space.Advances in Neural Information Processing Sys- tems, 33:15338–15349,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8119d5-a7cb-4654-81f8-785fa9f46316 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Penalized proximal policy optimization for safe rein- forcement learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ad360b-aae8-4b5f-9216-ebe30b0d1246 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9a9963-f392-473d-8c3f-0aba7a6b5bd4 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a380bf07-c3c3-4ec8-ae61-ac568aa8ef55 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ebec16-e3f4-49ef-88fc-7a321e6128bc · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Deep reinforcement learning for autonomous driving: A survey.IEEE Transactions on In- telligent Transportation Systems, 23(6):4909–4926,
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fae1680-ee90-429f-8558-5b3e8b73d0a9 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Augmented proximal policy op- timization for safe reinforcement learning
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a952f2ee-ef59-40ac-a35e-c9e192aad329 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Approximately optimal approximate reinforcement learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d0627e-ccd6-4b22-a505-a29acc84d432 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Routledge,
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b407b949-5a32-4bfd-bf24-2a9058e53444 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810483af-9f18-4f90-9d26-a9093c06bbea · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54d7a10-f8a2-4646-8b43-a66f2cf5d361 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Trustworthy artificial intelligence requirements in the autonomous driving domain.Computer, 56(2):29–39,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028dc12d-cf81-4fe9-9fea-c2e75bbae4e8 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Continuously Differentiable Exponential Linear Units
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67742a80-741a-421e-9d6b-3ffee382841c · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17694841-25c3-4539-942b-48e915c3b10d · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Natural policy gradient primal-dual method for constrained markov decision pro- cesses.Advances in Neural Information Processing Sys- tems, 33:8378–8390,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440be212-e38c-4ca0-a9a5-d1b7712cbf71 · outbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.