Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T15:58:27.229351Z
Paper Citation Record · LEDGER
As of 31 July 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:2509.12833.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T15:58:27.229351Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:12:07.125494Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-01T09:45:40.502853Z
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ecb2a73e-c959-408e-b241-acc9a4a73f1f · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Safe Exploration in Continuous Action Spaces
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 7c51347c-554b-4c1f-81fa-2dcdcf7e46ba · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Differentiable nonlinear model predictive control.arXiv preprint arXiv:2505.01353
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 73480cc4-5f45-4c84-aeea-4c1447930d07 · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Continuous control with deep reinforcement learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 6c4cdfa7-2043-4786-80f0-3cbda840d0c3 · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Asynchronous methods for deep reinforcement learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 4ed0a580-c6ab-4aa2-aee3-e294c5baba0a · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Fsnet: Feasibility-seeking neural network for constrained optimization with guarantees.arXiv preprint arXiv:2506.00362
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation e88586be-d83e-46cc-9911-6811e9b76bc3 · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d8845c75-f0be-4e56-9102-5ad1b74d35ab · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 35bfc727-ca6e-4a3a-96c3-95add16b4c5b · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Leveraging Analytic Gradients in Provably Safe Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 3231029e-e74c-4abe-b916-d316b705794f · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation a49a52ea-1d9c-4f57-a673-612badec4535 · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d563fcf4-d564-4be3-b07a-ce3c05cdc46c · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? ∫ ˜X ∫ U
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 38d69e1c-4123-4f09-90cf-ef54cebc6df4 · outbound
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? The environment has the state x= [ ϑ,˙ϑ ]T and the dynamics ˙x= ( ˙ϑ g ℓsin(ϑ) +1 mℓ2u ) ,(53) wheregis gravity andm,ℓare the mass and the length of the pendulum, respectively
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 0df37d30-5c17-4231-9f6b-c51e40a48a88 · inbound
Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 20ef1545-8009-4f72-abe6-15934da4e59f · inbound
Safe Online Learning via Smooth Safety-Structured Policy Composition Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.