Pith. sign in

Paper Citation Record · LEDGER

A Provable Approach for End-to-End Safe Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.21852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21852 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:26.255011Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5caa45fb-6d24-48eb-8d23-c96a11334fb7 · outbound

This paper cites Achiam, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Achiam, D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.889234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.727315Z digest=sha256:94ebf59bce7a3f903888042774d36480c848f8c2aa34a3a1b878cb38dd3c16d6

Observation 1ea45edf-8f6a-4412-922d-31cdef7e8a52 · outbound

This paper cites Alshiekh, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Alshiekh, R

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.708658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.857266Z digest=sha256:7d11aab925112cdff085d460d33bef84177a41bebf194727311dc7a975a022cf

Observation 4ef9eeff-cff6-48f8-a334-a5a74e082f37 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:37.456951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.994834Z digest=sha256:6655b3dcf32086e6e6e1d2dcaa23a36eb55c5e3de1c6c7cf4190873199b02c99

Observation 7a62189b-0e02-41e8-b30e-bfed14d055d6 · outbound

This paper cites Concrete Problems in AI Safety.

A Provable Approach for End-to-End Safe Reinforcement Learning Concrete Problems in AI Safety

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.092151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.092151Z digest=sha256:78f91265415bdd31066ad36c1c845d8aafe2b1625ba4dd6d90f6f8dfded648b8

Observation fc6ff8e2-20c2-4c3b-86b9-90bfad22a305 · outbound

This paper cites Berkenkamp, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Berkenkamp, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.287804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.255035Z digest=sha256:9ee3174b1eb30de52da02e8cb38272f3da16f2f101f70bb3994c8b095a63a29b

Observation 74ccda11-804f-4a36-9ef6-bd3f6e011539 · outbound

This paper cites Bhatnagar and K.

A Provable Approach for End-to-End Safe Reinforcement Learning Bhatnagar and K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.998638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.443200Z digest=sha256:9194830d7b334c9032d289db90fe4307f08b8bd2d69ff2fbcdfbf19d0f292c3a

Observation 5a2a7343-b76f-4fad-ba6e-a7991a8f7ccb · outbound

This paper cites Black, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Black, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.763546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.564755Z digest=sha256:1b24df6262a5ab0c0e5c9c4d4f5ba30a067f10f941292043c845f185cac26dae

Observation 32c8fa42-bdd5-451d-86ac-026819081ba9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.610803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.657269Z digest=sha256:7c694b69c5aab0f854f89a97315c2d8c9b4334db4d93bfe8a75661009680c5a4

Observation 8d0dea74-9e4d-4cc5-a7a2-0adfaa9148bc · outbound

This paper cites Brandfonbrener, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Brandfonbrener, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.444884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.882152Z digest=sha256:cd464d378bb6e4567ac57900382587efa3b26cb881f7e287628fb499f71f7396

Observation 4f6df005-2369-4bd7-a90c-d7466538cc85 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.211969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.964019Z digest=sha256:93ed7e0434a96630a34fb4945f949833024c317e830e111f35683fee816570ac

Observation c9dddd1d-31a7-4780-9147-a4281bd3b44e · outbound

This paper cites Cheng, G.

A Provable Approach for End-to-End Safe Reinforcement Learning Cheng, G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.005293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.094841Z digest=sha256:678b3efb3d9deac5071f257b475c5307c8d76e833ec6989de48e12119798f90c

Observation 6e2775c4-5441-4b23-a1c7-88659cd599c9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.833182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.171133Z digest=sha256:c3abea86370510679256335c8e26c0db546d6a9ec70dd0cb8e1683196d803add

Observation 674dfce5-6ac5-4f54-bc8a-68d633f7205e · outbound

This paper cites Da Costa, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Da Costa, M

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.294753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.294753Z digest=sha256:cca781ff6b0a8c8e1fa63d81eb0e464a7894ea41d5e805db29fce88c35a4f97e

Observation b13a745c-023b-40b8-8a72-08602e0cbedc · outbound

This paper cites Emmons, B.

A Provable Approach for End-to-End Safe Reinforcement Learning Emmons, B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.700611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.394376Z digest=sha256:09c26a6232642b175842e0767fa23eb903b1df154e36bb3d0757e0c90c6ea728

Observation 1e5e9932-f7f1-489a-bb75-3eba8f490110 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.525264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.491943Z digest=sha256:3e307e4101a9663b2512bc08e00120f94695988b269b79b0ed710dfeb3cce62c

Observation 092bc828-ae16-4493-8bed-b153b1723c04 · outbound

This paper cites Fujimoto, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Fujimoto, D

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.288966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.614956Z digest=sha256:b493008d2811b1968dc987d0750a01c898a01703fc1a3b91135f535cba996d49

Observation f5e404af-41b3-4494-966c-3c5cc3b0f525 · outbound

This paper cites Fulton and A.

A Provable Approach for End-to-End Safe Reinforcement Learning Fulton and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.105286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.692574Z digest=sha256:8a4a51cff80b0f90bcbc0f4e52003d50898c25570450391bd12e93062c0fbbed

Observation a7274ed4-4e90-4f4e-83fa-9173b709f53a · outbound

This paper cites Garcıa and F.

A Provable Approach for End-to-End Safe Reinforcement Learning Garcıa and F

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.901866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.790603Z digest=sha256:cdec11da78a62d00bb3b5004db06c7ba8b1509c86fc17e9ae3bf163de8afb8ed

Observation 97fcee7b-9fd3-4b4c-8a43-f7e7b001e6aa · outbound

This paper cites Gronauer.

A Provable Approach for End-to-End Safe Reinforcement Learning Gronauer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.699316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.937413Z digest=sha256:74496c8519ba617899b2c6bdd2954f4085f344e04563b8a5f0f2a95bd9f30d40

Observation c32c7df1-e8cf-4c9d-b28f-55a1f4bf4c8d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.500351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.049343Z digest=sha256:e7c232893b8d2c1c5bf0478c28926fc7bd25f7ba16873c247e9c4068d16da44d

Observation adbed44e-9a9b-4514-a306-d885ac1df103 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.159218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.159218Z digest=sha256:69ea9a6c21cfd70872d40e488e3d178e25bc7de02d5d1c79eb7b272ce2ee01d2

Observation 690ac445-dd81-405f-8f18-d5424fdf5081 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.344303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.245173Z digest=sha256:8345d29acffb21abbcc4958b588ef1b68bc27fb212a4edbdc95daa385f1af066

Observation dc55872e-60ae-495a-90c6-99f555156613 · outbound

This paper cites Hambly, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Hambly, R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.049121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.326144Z digest=sha256:7ad37197969250fe520288da79fb345f0c29777d97cd7f95cc2202ddcd1660c1

Observation 00fbdec3-2754-4655-acb8-94c756ce42f1 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.895147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.391642Z digest=sha256:5bc4be4222f79ebb6792011a29e37c91cf0607b84a608f1e1f6697c3be09250f

Observation 6897be8e-84ed-4d96-868a-a3899270c6c7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.732780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.504885Z digest=sha256:c0edde49760c5a0864262577d222948a78e2ef0b51378cb0a2fc852d3ac7c6c5

Observation 8e5fce55-3c99-4ee3-a2a8-9332af96a67c · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.570967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.637476Z digest=sha256:b1d78f04fa096b47dda58e179593d579720b7f93daca4ce582db15e04ff91df4

Observation cb9f6868-d166-4fd2-9fb6-db9ed9a6026c · outbound

This paper cites Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking.

A Provable Approach for End-to-End Safe Reinforcement Learning Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.701820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.701820Z digest=sha256:b73f1389c6c3c254f98a79fd1996b669f90dbeb525546717bfdde850f01a707f

Observation 02dac6d4-7dbd-4165-993c-83a9995dff40 · outbound

This paper cites Kumar, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Kumar, J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.369305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.786585Z digest=sha256:201f6b6e1692e9f3bb4fb467d1604790d1e9c375673023e282e6c1388cc1d521

Observation 0124c1ed-df2a-44cc-8b23-59c16ca05f04 · outbound

This paper cites Reward-Conditioned Policies.

A Provable Approach for End-to-End Safe Reinforcement Learning Reward-Conditioned Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.843583Z digest=sha256:1e4301976a76d511b265c234fca6fb85467d10146526799f63c0e23384615a61

Observation 139671e5-860b-468f-97ce-5e25eb919d11 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.098066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.975095Z digest=sha256:112be1dc5fd8418bede821e82f1ef7cb96458cb53e3a2e88dc7e76707fc73827

Observation febdd373-fd80-472d-9a82-065076405ea4 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.875494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.089264Z digest=sha256:012915f814703b0bf212b46cefec0f15a25253939b7511f6419ad00ecb654829

Observation 6ea14459-59f1-4823-b1e6-d0a9d8720cdb · outbound

This paper cites Levine, C.

A Provable Approach for End-to-End Safe Reinforcement Learning Levine, C

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.638209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.184252Z digest=sha256:586aae3a9756ee10b3e0bb267e34078cedc35128356c1f3237483439b86859c9

Observation dd2a1b61-3e23-49c2-97e9-7bb23ee939f8 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

A Provable Approach for End-to-End Safe Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.279554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.279554Z digest=sha256:f27bfec8cfadfb7c784336060e8565b2ce6d11dac1d015e43bb285cb8f45dd11

Observation dc0847ab-7a35-4193-904f-bde708311a34 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.392354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.333956Z digest=sha256:50a8f5cb63a923e6a46994fd9b97c83dfac48352d1561c06cfa1b46c941c113e

Observation 93140e8d-fce0-41ad-a807-5a00ac9cbac0 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.244747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.434878Z digest=sha256:8b5b675777c5a96815497318195fe0e6644e3785b008e8e3948177332b3926bc

Observation f0751b9c-ed33-46ac-bf04-49112ec2ab7d · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.525259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.525259Z digest=sha256:e96493e6bff3a803be8c60e88181975cd28813f48242b10e29430cfaa4b64bd3

Observation da4a3e35-5a01-478f-b769-9c6494246e5e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.017371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.654841Z digest=sha256:e39caea59382a74924a4276f5586ba6d9a85001a14b9d4c3f3d721d26c2b645f

Observation dc6ab268-0dca-444c-80f9-9dea7274a561 · outbound

This paper cites Ouyang, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Ouyang, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.804742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.724974Z digest=sha256:7677f78115dce93338c2da119138b4b58461a6fc47cbc785e8e7a670a0528e3e

Observation 10e93eee-0ed5-4ce2-8e51-32cbdd160ab2 · outbound

This paper cites Safe Policies for Reinforcement Learning via Primal-Dual Methods.

A Provable Approach for End-to-End Safe Reinforcement Learning Safe Policies for Reinforcement Learning via Primal-Dual Methods

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.002527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.813683Z digest=sha256:8396a475ffa6cd8a275281cdcba4fa28086ecd1aa2155b6ccffc1312b1f1ae10

Observation 4d25c2c2-a451-4962-96da-cd2e252c5317 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.872152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.872152Z digest=sha256:bab7a9bf0a774795612181153c408e672bf140785470365602533f6745f7d3ad

Observation 3871230a-2563-4852-b4bb-ec764f5f4fae · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:31.514817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.963485Z digest=sha256:064c272beec398d66c0c6356d875dc8d4d4947681a738254b9c10afb08b21d43

Observation 4ff7817e-ca91-4258-844d-b00f8440da8f · outbound

This paper cites Satija, P.

A Provable Approach for End-to-End Safe Reinforcement Learning Satija, P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.347590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.009443Z digest=sha256:028b211a29fda05f8d1d45e2bb6341878c4ab297ccb47cebfb44cd08abb99033

Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.086262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.086262Z digest=sha256:85c2d3cae91884548a7e2f07fc0c45b010a4bc472755ed9bbc332228ef15b0af

Observation add3ead9-d96b-4ea1-8ec3-d1b2e0eb72dc · outbound

This paper cites Sootla, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Sootla, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.171392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.163137Z digest=sha256:1f170b986992eb957c9d51803991693890422908aebbdfc33842936f55aa769d

Observation d66c4c7e-dc5a-4ada-9017-3768a69efe62 · outbound

This paper cites Srinivas, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Srinivas, A

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.214658Z digest=sha256:cced79b08f55966038c58c228d0a702b82aa4e3e4a4084a694fb371b97c187b0

Observation fc5c38cc-235b-47fc-a44a-eabe153027c2 · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.270417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.270417Z digest=sha256:475b007d8fdaaf0fc79d33df4af2fd6958c6afc2915b30bf5e2aa3b1c3e04ffa

Observation 926187b4-167e-4697-8738-b0f9704de70d · outbound

This paper cites Stooke, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Stooke, J

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.791525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.345683Z digest=sha256:cbe6eebb62763dc48df488592f3c1a2e7a02b2dd9b927e590929446d3c1f6a7e

Observation d7812556-4fad-48c9-89e9-0720b8ceea70 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.584763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.514869Z digest=sha256:1a59a854169c53df75c5f7a3b827ca1559cb09f763020d7e43b072424ae740bf

Observation 3469f3d0-0a5b-49e6-8e86-57cabe6f424d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.442925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.636311Z digest=sha256:d644d40baf724311da179b5fb7a067939672fb2002c982f1e82f555fd42479ad

Observation 47120573-a122-4411-92b6-ec129f9b34d5 · outbound

This paper cites Turchetta, F.

A Provable Approach for End-to-End Safe Reinforcement Learning Turchetta, F

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.277183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.735023Z digest=sha256:26f2ac63931d5333d9f7427c3d745df3446c47ce84a6940791518596bf5c64f0

Observation b1411c21-861b-4134-81b4-67a501ca461e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.804655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.804655Z digest=sha256:d111c8539887f5b76d13425443f903c99f2ccd8856dd2fc69cec8e26fe885e24

Observation d32a2ba4-93e9-441a-a0b4-beb49d816e7d · outbound

This paper cites Wachi and Y.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi and Y

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.104278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.914873Z digest=sha256:85057056aa5e9598910dec058502e90e136f10efc70010cd209258260bdaa3ef

Observation a0062dd8-7c4c-476a-aea2-d501f869a7db · outbound

This paper cites Wachi, W.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.874805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.018298Z digest=sha256:fec7144491619b29c54f32fb696618c5324742b434d259a070e24c601b0d6c6c

Observation 327c2ee6-cc89-40f5-8722-4dfe318abc7f · outbound

This paper cites Wachi, X.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, X

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.713700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.234827Z digest=sha256:29660e19b139be22acc914d708c6bd710de5096f2c36263ffe530b248f218ca0

Observation 3cfc6953-bd2c-471a-80eb-5e8de9f2eb1b · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.524828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.375135Z digest=sha256:e2460de910cb0ac191c3afee9ae4ad3995984973727851aec4a8a5c7d3036ff4

Observation 4400f6c7-0285-4bb1-9c33-50dd9d5be1a7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.299540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.644735Z digest=sha256:b689a61ca8f7ddc9e376482b5b25f5943061af8f2fa51ff48ac345a2c0a35751

Observation 23b55a16-b18d-49ce-a66f-b55830d03c47 · outbound

This paper cites Yang and M.

A Provable Approach for End-to-End Safe Reinforcement Learning Yang and M

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.074292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.815177Z digest=sha256:bef4d87dd9c2f71d7b207e78f354c7731663f1f8cbc515da784ef37c475209b8

Observation 701d1300-f321-4120-aba9-54be7671f6b7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.892426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.934915Z digest=sha256:63107786ae7d24f64da6529e4fd39307d8bc5c5f4dddcc48500ea1e8895479cc

Observation 92c4b7fb-3b1e-4c95-9a80-e8f6a8c5dfe6 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.742128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.009155Z digest=sha256:fa30b41e460d9f2010f44d1f50a82b749bb18b56a5e50e686ae4d93e504b323e

Observation d7a6d77d-4f82-41b8-9618-8e2501cd0dcc · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.579001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.125071Z digest=sha256:d8f73358597f453608d0477b9db02e3ae92d65e1347101d6e31aab3809a15f73

Observation 6362a78c-7d61-49b9-95da-5e8cb5b3a191 · outbound

This paper cites Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d.

A Provable Approach for End-to-End Safe Reinforcement Learning Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:28.467897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.255011Z digest=sha256:11d652705c40d535e717e6e98642dfbaaf5f1934e7c7e93117d3badcbce46a62

Pith citing papers

No inbound Pith citation observations are available.