Pith. sign in

Paper Citation Record · LEDGER

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.08925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08925 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T05:46:03.704125Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3218436a-3f43-4a88-a9ef-215fee392dff · outbound

This paper cites Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:669aa2374d4f5f1f9d11c9b9f38743a7e292d274413f0dd0a0d9ee55694e8b1b

Observation 374f5067-d0b1-4270-8333-5a612d899c26 · outbound

This paper cites ISBN 9781605585161.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions ISBN 9781605585161

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:522bb375352c843901d0217a66d085159bd5642b7655633258604e7a09cadc24

Observation 1ef3ff0c-df5d-4e33-82bc-da561075d2e6 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Safe Exploration in Continuous Action Spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:3ee6537b872f7b076c7816043cf138b5be84a2be8e7e252d51e27bc3de21f84f

Observation 6ff456a5-5f81-491e-a958-8b54531a3919 · outbound

This paper cites Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening

Reference 4

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.740269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:c2ad86ea8a165a31d277ea0c16f08e64820c59023731acb0dd9b7897a68fa996

Observation f32a5ce2-e4b7-44ed-9bd1-c20276b28357 · outbound

This paper cites Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and Jo ao G.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and Jo ao G

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:4c5e5d6275bae90dfab85c7a3dccd763b57940241c36dfbf0589148745f58955

Observation 23c698ac-9b57-4aad-8d25-47142d5a8b12 · outbound

This paper cites Choi, Michael Janner, Claire J.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Choi, Michael Janner, Claire J

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:b28e221b31333503ff50939612a50417d068a6cca470a05759c40a36ac7b3ab7

Observation 561d899a-c361-4a7d-a6b1-30e1e74cffbe · outbound

This paper cites Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:7d658556a2ea017ef70d77409c5bf9277a0f99d4b1aa3d0678740c4bec04b3a0

Observation 66ea9a09-f858-46da-b009-5dc576c9b0d3 · outbound

This paper cites Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-13T05:49:23.735751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:b904297790cddea9c2b5d4e8c743cb8a6c33f54ffe100808c4f483776eb44fe7

Observation 358108df-7596-4570-8fd8-9900570da405 · outbound

This paper cites Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:76da5e461f96d5c5db0caec3350f68b23fc9705263a770c63ae79f1512d732ce

Observation dccd144d-696a-4a93-96b5-cab7972b1249 · outbound

This paper cites Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Edward Lee, Jie Tan, and Sergey Levine.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Edward Lee, Jie Tan, and Sergey Levine

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:871d920182962daf2c3b322f2765c5f603a74a18ef72478d2425a19d141fd8ec

Observation 2f8ac172-9061-47f4-a0bf-f0aaad2a9e55 · outbound

This paper cites Xun Pua and Majid Khadiv.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xun Pua and Majid Khadiv

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:959771c56d65991a90f8ae0ee6f698d1b8d3637771219fc71aec35fdceffd479

Observation 074112a7-aa09-4956-887b-1a17c04d10ff · outbound

This paper cites 2024.10769799.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions 2024.10769799

Reference 12

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.732396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:cb27ec91e0c64b381daf8003c69c0c7769d827a28a6aa280e7a4da790329a7e7

Observation 61f99978-a860-4a60-828d-9d1ff0eeb965 · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:1ca5870d1c145ea9bee7876916116d536c456e012a8dde8a2c173713d9f74265

Observation 7f1f459e-8eb2-4ee1-a0ac-7e6defe77733 · outbound

This paper cites Trial without error: Towards safe reinforcement learning via human intervention.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Trial without error: Towards safe reinforcement learning via human intervention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:fdc39acfa7cb9bde45b4d420de7cb910bafb65cd6b4ebce6d471aa376b607c3f

Observation 258d8482-dc10-4668-9ef7-1a446d1fb73c · outbound

This paper cites Proximal Policy Optimization Algorithms.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:29de62ecf76a911c1127b6c625ec461f627573e70b9f10ae7e9cdf593e9801a9

Observation 71ae7c4e-07b1-4d91-9455-667a6c5dbe44 · outbound

This paper cites Smith, J.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Smith, J

Reference 16

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.761511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:1f8bd44ab5235b0912bef14d4b621878847046851a0ed3c5ae5bab8bcec9c68b

Observation 47bd04e3-9676-486c-956f-e380a6bf15f0 · outbound

This paper cites Learning to be Safe: Deep RL with a Safety Critic.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to be Safe: Deep RL with a Safety Critic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:b9b83081fe8d9f7b7ef2c17c9e4f79e234b3efa5d20c9c00743473a77ce95c78

Observation 6fbf0afc-f137-487a-8e8e-33a50aeca87e · outbound

This paper cites Sim-to-real: Learning agile locomotion for quadruped robots.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sim-to-real: Learning agile locomotion for quadruped robots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:04e9b31e89bc9f5d344ca8228f2e414b87303bc3bc3e6bf6b426f120c846a7ee

Observation a91ff146-8d42-498a-bb9e-1893a6f293b2 · outbound

This paper cites Reward Constrained Policy Optimization.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Reward Constrained Policy Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:967a0b8f915503984e7628603ce2915b17f5664de9605e72cae67c0c152a4b38

Observation 56cdc0bd-9fc4-48a2-a58c-b8b6cc9060b6 · outbound

This paper cites Emanuel Todorov, Tom Erez, and Yuval Tassa.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Emanuel Todorov, Tom Erez, and Yuval Tassa

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:29c9bbfb45e8eb16582b201ea383fc59f71d6084afd00163bd3677b6dfa3cc6b

Observation 06645f40-c119-4ad3-a0dd-57568a859669 · outbound

This paper cites Mark Towers, Jordan K.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mark Towers, Jordan K

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:0b554fdde754389f85786f7bb1005ec13c0214095637ec8d97fbb4004a560b9f

Observation ffc01a51-20ed-4a4c-bfd2-176bcd8158ed · outbound

This paper cites URLhttps://doi.org/10.24963/ijcai.2024/913.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2024/913

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:cddfb543ba131e2b8462e8fef2fb80b1f253c30f13478b7f752f73c36d3c2f7a

Observation f6355c3b-f2b4-4c47-90bf-b0fef8aed300 · outbound

This paper cites Kevin Zakka, Yuval Tassa, and MuJoCo Menagerie Contributors.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Kevin Zakka, Yuval Tassa, and MuJoCo Menagerie Contributors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:194d80d0e9085d75b0e60bccab709d19691844863c7901c7dfb718a88b4b01ce

Observation 4da7ee70-3e09-4f31-b7a0-41fbf86662a1 · outbound

This paper cites URLhttps://doi.org/10.24963/ijcai.2023/763.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2023/763

Reference 24

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.749250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:c046e602e9cf98d8fbdc8bc3c7e542fbc935d06c0e18ff2d8efb8e1c2331b4b6

Observation 2fe6394f-3268-4bcf-9cc2-11947860fb17 · outbound

This paper cites Clipping bounds variance downstream of the singularity rather than removing it.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Clipping bounds variance downstream of the singularity rather than removing it

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9db005a546d4483baa2add84f8b3a90322e7b37abe47c759428d5f980bf2b030

Observation 30d9b33d-1e35-4927-a0c1-db54f7ecacd9 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:5be2d9b59763a08e68f8ce86484a12217a453809499b679d43522cb81af672d5

Observation 7879240e-bc99-45df-b474-a43106d4cbba · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9ca4a65f6eb17f5339adbcdedbe572f1758fc9b68c4aa1a5caf8754b50743bcc

Observation 281778e4-fe75-4cbe-b09f-56f2c4bb271f · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:2d3e6c3c73d01f6ba6cbd4fa99510a0e8ce5bd0c7946051dadd8fedc0409da62

Observation 9722ad14-7b7e-44dd-89cf-5a72673eead8 · outbound

This paper cites CPO and PPO-Lagrangian additionally carry the constraint hyperparameters their objective requires (Section C.3).

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions CPO and PPO-Lagrangian additionally carry the constraint hyperparameters their objective requires (Section C.3)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:32b7992e48cefb0d9304c9942e7e204d0c82673709e1f3eb0e57569497dc6a79

Observation 36000acf-9365-4be8-a4b0-19692dbfcb5d · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:1d527c31a014ca03b5810abf4d7d51d5146d298b2a53a743bda65ddf6f067bae

Observation 5209fcc6-7cba-41d5-8ac8-b87faf89dc3e · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:8611c2ec968a300b9ce38804a37aa1954010581f94de3515aebca2f3bd44a6eb

Observation ea2c3ada-c1ea-42b8-a5d1-ffe5a9adc628 · outbound

This paper cites global”). An alternative “per-segment.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions global”). An alternative “per-segment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:3f23b9363856f37a83a9de8f73b4c4eb3167d11e099eb32113db0dab295948e4

Observation b2ce74f8-3002-4757-8960-f03e60a71d89 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:7a766d4e19930f88c8d0626430fe683f7dd3a2b754735bd910f136eeeba6c0c0

Observation a347c806-4456-4019-bf06-729146d76549 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:704b8e39434e8fcd69fb6d66219af2df136dfdd047fca99c99d656d46a0be353

Pith citing papers

No inbound Pith citation observations are available.