Pith. sign in

Paper Citation Record · LEDGER

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

As of 31 July 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:2509.12833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.12833 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T15:58:27.229351Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:12:07.125494Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T09:45:40.502853Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecb2a73e-c959-408e-b241-acc9a4a73f1f · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Safe Exploration in Continuous Action Spaces

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:01:34.683098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:59466974ae22b3dc2f7860162ec0e3f35ef811336c169a0e8cf99644433bfb35

Observation 7c51347c-554b-4c1f-81fa-2dcdcf7e46ba · outbound

This paper cites Differentiable nonlinear model predictive control.arXiv preprint arXiv:2505.01353.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Differentiable nonlinear model predictive control.arXiv preprint arXiv:2505.01353

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:01:34.669220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:34da6bc32802d6b990b93550debf4cf1fdb2603cb20e402693b5d063b771ae8e

Observation 73480cc4-5f45-4c84-aeea-4c1447930d07 · outbound

This paper cites Continuous control with deep reinforcement learning.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Continuous control with deep reinforcement learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:01:34.691712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:33d8a01010db86d84de043a93ac12d833db29ec323d1ab1b0beff46310a8966e

Observation 6c4cdfa7-2043-4786-80f0-3cbda840d0c3 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Asynchronous methods for deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:02:42.413773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:95fc2cfb6dd4902f673ce7a9539ec94f2df87cadba60859858e1586040a23d8f

Observation 4ed0a580-c6ab-4aa2-aee3-e294c5baba0a · outbound

This paper cites Fsnet: Feasibility-seeking neural network for constrained optimization with guarantees.arXiv preprint arXiv:2506.00362.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Fsnet: Feasibility-seeking neural network for constrained optimization with guarantees.arXiv preprint arXiv:2506.00362

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:01:34.697150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:25ee3cf5cfa62dd703285f8708da48c22239aa8a7b757ecb0d704fe7c758e297

Observation e88586be-d83e-46cc-9911-6811e9b76bc3 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:01:34.687253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:485f754bce1c27c37431c73d1e16c2ecfb7548c676f05e74a5afb052dbc24534

Observation d8845c75-f0be-4e56-9102-5ad1b74d35ab · outbound

This paper cites Proximal Policy Optimization Algorithms.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Proximal Policy Optimization Algorithms

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:01:34.673750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:978c63f4d34c77740e9e92e4c084c3fcd7bdda296468eb5849f7f670e470b207

Observation 35bfc727-ca6e-4a3a-96c3-95add16b4c5b · outbound

This paper cites Leveraging Analytic Gradients in Provably Safe Reinforcement Learning.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Leveraging Analytic Gradients in Provably Safe Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:01:34.678415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:4e5028ca8407c001884167118750a35575f9ff1991b1f43421a60bc38c0f6223

Observation 3231029e-e74c-4abe-b916-d316b705794f · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-18T16:02:42.416405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:5d02e7ad43e2f4cbdbb7f90e940e15f94cce5681e3ed414ceb196e55db136bec

Observation a49a52ea-1d9c-4f57-a673-612badec4535 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-18T16:02:42.403860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:9803f39c32993df9e1ec0cc61a322507a86163c9bc095a6e5297ec1fd8b68115

Observation d563fcf4-d564-4be3-b07a-ce3c05cdc46c · outbound

This paper cites ∫ ˜X ∫ U.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? ∫ ˜X ∫ U

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:02:42.410567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:1ba9c3d57a4f6d84d2e88ad84719df9ef5a315291821c27240f86b08631fc0b2

Observation 38d69e1c-4123-4f09-90cf-ef54cebc6df4 · outbound

This paper cites The environment has the state x= [ ϑ,˙ϑ ]T and the dynamics ˙x= ( ˙ϑ g ℓsin(ϑ) +1 mℓ2u ) ,(53) wheregis gravity andm,ℓare the mass and the length of the pendulum, respectively.

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment? The environment has the state x= [ ϑ,˙ϑ ]T and the dynamics ˙x= ( ˙ϑ g ℓsin(ϑ) +1 mℓ2u ) ,(53) wheregis gravity andm,ℓare the mass and the length of the pendulum, respectively

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:02:42.407647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T15:58:27.229351Z digest=sha256:0eb5b06913a5ba8e0248175784f7e4470208934b6c225caf94a12c5e5e905163

Pith citing papers

Observation 0df37d30-5c17-4231-9f6b-c51e40a48a88 · inbound

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty cites this paper.

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:36:39.984535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T16:34:33.155628Z digest=sha256:2f08c3b3e8d490bd2d3c3f5f43caa5da5579362b33efda266bddb1da440ce71d

Observation 20ef1545-8009-4f72-abe6-15934da4e59f · inbound

Safe Online Learning via Smooth Safety-Structured Policy Composition cites this paper.

Safe Online Learning via Smooth Safety-Structured Policy Composition Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:40.504080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:9522b5928347109ea147d09ad7850ddbb6ed41f2e758c67ce60fdf22a5b96995