Pith. sign in

Paper Citation Record · LEDGER

A Provable Approach for End-to-End Safe Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.21852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21852 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:26.255011Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5caa45fb-6d24-48eb-8d23-c96a11334fb7 · outbound

This paper cites Achiam, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Achiam, D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.889234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:19.727315Z digest=sha256:fd63301cda0c62a5bed8fbc105015f3262150282ff4d4bd045a3434ee83dc2da

Observation 1ea45edf-8f6a-4412-922d-31cdef7e8a52 · outbound

This paper cites Alshiekh, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Alshiekh, R

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.708658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:19.857266Z digest=sha256:af08c3732b1891fe11210ffd9cc4f210556ad69e39f7464ab49cc1513096361f

Observation 4ef9eeff-cff6-48f8-a334-a5a74e082f37 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:37.456951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:19.994834Z digest=sha256:b99d50021d485ad64b2bba0da47358cf4238a941a8a63fa26c77d5bce6ac0e6e

Observation 7a62189b-0e02-41e8-b30e-bfed14d055d6 · outbound

This paper cites Concrete Problems in AI Safety.

A Provable Approach for End-to-End Safe Reinforcement Learning Concrete Problems in AI Safety

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.092151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.092151Z digest=sha256:723f6e3d242989dafd8e907808b5c01fb0305daa69d82b14b486e03f180261d6

Observation fc6ff8e2-20c2-4c3b-86b9-90bfad22a305 · outbound

This paper cites Berkenkamp, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Berkenkamp, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.287804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.255035Z digest=sha256:fcc94170f75b95e585886d310a1acfc58b0a00fc6ef5e46c49045d32b7b97680

Observation 74ccda11-804f-4a36-9ef6-bd3f6e011539 · outbound

This paper cites Bhatnagar and K.

A Provable Approach for End-to-End Safe Reinforcement Learning Bhatnagar and K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.998638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.443200Z digest=sha256:1cf296876a9e69234e31d8890ee4ce34f01996908e4d711550c4ca9194b6ebfa

Observation 5a2a7343-b76f-4fad-ba6e-a7991a8f7ccb · outbound

This paper cites Black, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Black, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.763546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.564755Z digest=sha256:81cfe66cffaee5bbb528e8b1b3e212a2ac46d101d311435a8ad3ceeeb7a5982b

Observation 32c8fa42-bdd5-451d-86ac-026819081ba9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.610803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.657269Z digest=sha256:5875fc7c7d246c4466a7e0e952c5d9ea4a0ad09f44b50146904c147299397cc6

Observation 8d0dea74-9e4d-4cc5-a7a2-0adfaa9148bc · outbound

This paper cites Brandfonbrener, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Brandfonbrener, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.444884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.882152Z digest=sha256:1974f034f5a6f2b79b154834f8439b51c058cd02510f1aaac7d6aa0e0024c7b5

Observation 4f6df005-2369-4bd7-a90c-d7466538cc85 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.211969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:20.964019Z digest=sha256:6365b0d612256cb8d5426340b1c9a073084aa328b925694a99ec659dc903039f

Observation c9dddd1d-31a7-4780-9147-a4281bd3b44e · outbound

This paper cites Cheng, G.

A Provable Approach for End-to-End Safe Reinforcement Learning Cheng, G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.005293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.094841Z digest=sha256:eb5eb0ad95419f978f4fa5fb94a16efa6a4d7c266387cfdba1b2bf2ad3168d5a

Observation 6e2775c4-5441-4b23-a1c7-88659cd599c9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.833182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.171133Z digest=sha256:ad036802d8a83d4c7aa398d53c20b34aa2d26ae4dababb173552633d2030ea67

Observation 674dfce5-6ac5-4f54-bc8a-68d633f7205e · outbound

This paper cites Da Costa, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Da Costa, M

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.294753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.294753Z digest=sha256:214f53ed738c60c49db3286ccfae63b41f19e02590453fddbf2438c9e3d09787

Observation b13a745c-023b-40b8-8a72-08602e0cbedc · outbound

This paper cites Emmons, B.

A Provable Approach for End-to-End Safe Reinforcement Learning Emmons, B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.700611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.394376Z digest=sha256:903fe484923fcbda9d20abd458fa3dbc79951cbc110baa795e38efe82a97cd7d

Observation 1e5e9932-f7f1-489a-bb75-3eba8f490110 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.525264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.491943Z digest=sha256:5e8342a49fb316d0282e793b269c6bd61f1b4e05b8a4607bb991ff5cbd618db4

Observation 092bc828-ae16-4493-8bed-b153b1723c04 · outbound

This paper cites Fujimoto, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Fujimoto, D

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.288966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.614956Z digest=sha256:871ecabcb4c1827e37c61da0e0690b203cbd989fe8aecf4d98c122313deb4ba4

Observation f5e404af-41b3-4494-966c-3c5cc3b0f525 · outbound

This paper cites Fulton and A.

A Provable Approach for End-to-End Safe Reinforcement Learning Fulton and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.105286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.692574Z digest=sha256:f015c0b4a5ef6a158fb5cdf727f3b789eebe6cc7efeccff9ecd55feb3845ee66

Observation a7274ed4-4e90-4f4e-83fa-9173b709f53a · outbound

This paper cites Garcıa and F.

A Provable Approach for End-to-End Safe Reinforcement Learning Garcıa and F

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.901866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.790603Z digest=sha256:b7bb987e6ef845ae5ae7c257eb9f1f17f3872e7c292f2a4ef3898b8e13fe4c68

Observation 97fcee7b-9fd3-4b4c-8a43-f7e7b001e6aa · outbound

This paper cites Gronauer.

A Provable Approach for End-to-End Safe Reinforcement Learning Gronauer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.699316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:21.937413Z digest=sha256:e2194618df7b64d5a1a5b560282216d9f7765e3dfba267de309659b9e4a11a3d

Observation c32c7df1-e8cf-4c9d-b28f-55a1f4bf4c8d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.500351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.049343Z digest=sha256:ea3ff7684ee28871593a0abb193b7afcef2e0ba838b8506e82a9296a26764f69

Observation adbed44e-9a9b-4514-a306-d885ac1df103 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.159218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.159218Z digest=sha256:0b2f04f353fcfc90c02cbb6ea0c879d74bc822d84692555d2366f30bb1907514

Observation 690ac445-dd81-405f-8f18-d5424fdf5081 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.344303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.245173Z digest=sha256:ebee629327b159b19fae43bb12deb593e9e7e29988fa5534434ba5b14899bf01

Observation dc55872e-60ae-495a-90c6-99f555156613 · outbound

This paper cites Hambly, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Hambly, R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.049121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.326144Z digest=sha256:1af5e4eb6f59d8b3e959a6a9f29925ce1ac3f6aab2d1829c08b90552718d6b50

Observation 00fbdec3-2754-4655-acb8-94c756ce42f1 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.895147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.391642Z digest=sha256:18bb76822570557e38e75abe5dc343e1984901dfc82c40a1a3c0fc5b5b4831c5

Observation 6897be8e-84ed-4d96-868a-a3899270c6c7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.732780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.504885Z digest=sha256:1df5724afdf01f98fdf02404312f623175051c4204668cc2c8d4a849eaf6ab04

Observation 8e5fce55-3c99-4ee3-a2a8-9332af96a67c · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.570967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.637476Z digest=sha256:aeceb52628dd42919aa1291c0bbf5cba768bdc5b2ab9cd4ae9dc46b21fe8baca

Observation cb9f6868-d166-4fd2-9fb6-db9ed9a6026c · outbound

This paper cites Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking.

A Provable Approach for End-to-End Safe Reinforcement Learning Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.701820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.701820Z digest=sha256:c5c67db7bef4acacb4ad39bb1593cc99dadf0b7a1771d2153a7c6a20209e7b62

Observation 02dac6d4-7dbd-4165-993c-83a9995dff40 · outbound

This paper cites Kumar, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Kumar, J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.369305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.786585Z digest=sha256:0c1aa363dd32946b7f4f8c1a181cdbd4c4d4119b988cf737565e0ecef5babcdd

Observation 0124c1ed-df2a-44cc-8b23-59c16ca05f04 · outbound

This paper cites Reward-Conditioned Policies.

A Provable Approach for End-to-End Safe Reinforcement Learning Reward-Conditioned Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.843583Z digest=sha256:c4a3d63f7a0cbf6426a62bdb7128a5243c67c422613bf14adeca8d7677c76027

Observation 139671e5-860b-468f-97ce-5e25eb919d11 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.098066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:22.975095Z digest=sha256:2368ae751a9dc241fcc1ff9e836439cfd6fca3c3226a3b3f8ecf2ecdbe7239be

Observation febdd373-fd80-472d-9a82-065076405ea4 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.875494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.089264Z digest=sha256:3e12507e6739301af805dff7bf303188b739305bb051fa9289b7192305492976

Observation 6ea14459-59f1-4823-b1e6-d0a9d8720cdb · outbound

This paper cites Levine, C.

A Provable Approach for End-to-End Safe Reinforcement Learning Levine, C

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.638209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.184252Z digest=sha256:2cb77e84267bb60ecaf295ded33b7a7966fe85f32b25b7616a11d7a802a0624d

Observation dd2a1b61-3e23-49c2-97e9-7bb23ee939f8 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

A Provable Approach for End-to-End Safe Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.279554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.279554Z digest=sha256:419cfefe0c40cfff9469376eaf1444ae909c616bae952c0e1cf50d4eab5af343

Observation dc0847ab-7a35-4193-904f-bde708311a34 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.392354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.333956Z digest=sha256:177f65bf4a3111609611dac6d06c688bd0273a29ff75f67c9926af0033b70df3

Observation 93140e8d-fce0-41ad-a807-5a00ac9cbac0 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.244747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.434878Z digest=sha256:1dd787809f6b7ad7c4338385725aa2fac0af73bc0738c099560c59cda65a6382

Observation f0751b9c-ed33-46ac-bf04-49112ec2ab7d · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.525259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.525259Z digest=sha256:abeb87d08fec0732a0db41c953b266f1f783665fe8fa47dac02c93e5378c2dc4

Observation da4a3e35-5a01-478f-b769-9c6494246e5e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.017371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.654841Z digest=sha256:5e993a3d2ff7da47aa01e2c944c4799d858d1d101df4683546342cbc3bc848c7

Observation dc6ab268-0dca-444c-80f9-9dea7274a561 · outbound

This paper cites Ouyang, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Ouyang, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.804742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.724974Z digest=sha256:139208472f44b94d2aae4cdd18a9a3f9f1a23a82da95996c1cefb31820548ead

Observation 10e93eee-0ed5-4ce2-8e51-32cbdd160ab2 · outbound

This paper cites Safe Policies for Reinforcement Learning via Primal-Dual Methods.

A Provable Approach for End-to-End Safe Reinforcement Learning Safe Policies for Reinforcement Learning via Primal-Dual Methods

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.002527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.813683Z digest=sha256:8db11b05c0e9aaf4e5aed69b7345f6ef69940e612e3fb53857371d19b7e74d0b

Observation 4d25c2c2-a451-4962-96da-cd2e252c5317 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.872152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.872152Z digest=sha256:835ec48b5271d5639c60a498529d11bd23c7d9d0838d80b0076a5eeeefaee342

Observation 3871230a-2563-4852-b4bb-ec764f5f4fae · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:31.514817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:23.963485Z digest=sha256:50c0d944d2e9e28b8e8261766e5d669851eca81ce73ce32acf7e630a4d02764a

Observation 4ff7817e-ca91-4258-844d-b00f8440da8f · outbound

This paper cites Satija, P.

A Provable Approach for End-to-End Safe Reinforcement Learning Satija, P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.347590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.009443Z digest=sha256:8b59d9a069cb13fcafaf08ff05dfba3643e283ff3b07c1fac6afa34660f945f4

Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.086262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.086262Z digest=sha256:34032d3e8e04ad3036bf0047fb40d074fa44185c3f47a19a4073155b545ce6f1

Observation add3ead9-d96b-4ea1-8ec3-d1b2e0eb72dc · outbound

This paper cites Sootla, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Sootla, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.171392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.163137Z digest=sha256:ca17a5c82bbd1ee7a44324b98d0c509fd4d963a736f39784b03ca87637b84f39

Observation d66c4c7e-dc5a-4ada-9017-3768a69efe62 · outbound

This paper cites Srinivas, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Srinivas, A

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.214658Z digest=sha256:fbd13add28b7cc9d58344e7c7e0ca6f3799d73b345a8f4bcd515661635635685

Observation fc5c38cc-235b-47fc-a44a-eabe153027c2 · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.270417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.270417Z digest=sha256:125d734320bff5c7ec3717bc039b847570ba643fe98b579153c6500fb0ca56f8

Observation 926187b4-167e-4697-8738-b0f9704de70d · outbound

This paper cites Stooke, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Stooke, J

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.791525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.345683Z digest=sha256:a90105931ddea5ffba414463a00c60c06ba727276112bb222b178ca9c5823f0f

Observation d7812556-4fad-48c9-89e9-0720b8ceea70 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.584763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.514869Z digest=sha256:c9ce6f42cb2b7150e5ba44ed27164a692a2543fa8331015f62e8ff101e79e465

Observation 3469f3d0-0a5b-49e6-8e86-57cabe6f424d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.442925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.636311Z digest=sha256:d042acbb456de341d39861f5b90a14a683b8da89a9bbc66093007f7983bbfc9c

Observation 47120573-a122-4411-92b6-ec129f9b34d5 · outbound

This paper cites Turchetta, F.

A Provable Approach for End-to-End Safe Reinforcement Learning Turchetta, F

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.277183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.735023Z digest=sha256:d0b14fbee582f8dace76971c535e617f72560952ded8d8326202c45ce6935e10

Observation b1411c21-861b-4134-81b4-67a501ca461e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.804655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.804655Z digest=sha256:ca10e398e71e13f2533a01782019651fdaf20f5c90654f4253097fb486352706

Observation d32a2ba4-93e9-441a-a0b4-beb49d816e7d · outbound

This paper cites Wachi and Y.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi and Y

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.104278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:24.914873Z digest=sha256:1ab5841efc6bcf4817d172f5ceedef43af0f067ed1d53a6b5fe239eab938a00c

Observation a0062dd8-7c4c-476a-aea2-d501f869a7db · outbound

This paper cites Wachi, W.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.874805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.018298Z digest=sha256:35d357df6f73cd2552a5fa614acc74188ef28e3fb983ada187f364280b951417

Observation 327c2ee6-cc89-40f5-8722-4dfe318abc7f · outbound

This paper cites Wachi, X.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, X

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.713700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.234827Z digest=sha256:573de5f4f0eee9dce4946a26b70be98ba3d5cfd6a298dbfe32c5e9b82b4f7d07

Observation 3cfc6953-bd2c-471a-80eb-5e8de9f2eb1b · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.524828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.375135Z digest=sha256:5709738815a8c722a9148275a9f7a752ec8887e45b9e9ee56f47f573dd1ab4d8

Observation 4400f6c7-0285-4bb1-9c33-50dd9d5be1a7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.299540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.644735Z digest=sha256:0c419c3ed502bd2f7931a962d2945e3cdaaa6fe2ddf3b4da9ae8c16fde36456f

Observation 23b55a16-b18d-49ce-a66f-b55830d03c47 · outbound

This paper cites Yang and M.

A Provable Approach for End-to-End Safe Reinforcement Learning Yang and M

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.074292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.815177Z digest=sha256:33f55fada2b852d61174e79bdae8ab661fd746f1f95a087340144bb48dd1c841

Observation 701d1300-f321-4120-aba9-54be7671f6b7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.892426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:25.934915Z digest=sha256:72c660bace4cd9edd068039e859116b0abc7937c1b3cac800480f29b79a93a76

Observation 92c4b7fb-3b1e-4c95-9a80-e8f6a8c5dfe6 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.742128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:26.009155Z digest=sha256:304d95e92f33ce3e245f6ab9f0757d09630eefa7b5c3cd25562532dcaf80484b

Observation d7a6d77d-4f82-41b8-9618-8e2501cd0dcc · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.579001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:26.125071Z digest=sha256:a14bdc35a3557881df144e98975b3e59f54a1997a815d35dc597f90538a61681

Observation 6362a78c-7d61-49b9-95da-5e8cb5b3a191 · outbound

This paper cites Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d.

A Provable Approach for End-to-End Safe Reinforcement Learning Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:28.467897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:30:26.255011Z digest=sha256:8ea66d5670028809091a1d378bb75005136757c298c89479b581f460566cd1d8

Pith citing papers

No inbound Pith citation observations are available.