Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T05:46:03.704125Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.08925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T05:46:03.704125Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3218436a-3f43-4a88-a9ef-215fee392dff · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374f5067-d0b1-4270-8333-5a612d899c26 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions ISBN 9781605585161
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef3ff0c-df5d-4e33-82bc-da561075d2e6 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Safe Exploration in Continuous Action Spaces
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff456a5-5f81-491e-a958-8b54531a3919 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f32a5ce2-e4b7-44ed-9bd1-c20276b28357 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and Jo ao G
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c698ac-9b57-4aad-8d25-47142d5a8b12 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Choi, Michael Janner, Claire J
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561d899a-c361-4a7d-a6b1-30e1e74cffbe · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ea9a09-f858-46da-b009-5dc576c9b0d3 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 358108df-7596-4570-8fd8-9900570da405 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dccd144d-696a-4a93-96b5-cab7972b1249 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Edward Lee, Jie Tan, and Sergey Levine
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8ac172-9061-47f4-a0bf-f0aaad2a9e55 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xun Pua and Majid Khadiv
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074112a7-aa09-4956-887b-1a17c04d10ff · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions 2024.10769799
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61f99978-a860-4a60-828d-9d1ff0eeb965 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to walk in minutes using massively parallel deep reinforcement learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1f459e-8eb2-4ee1-a0ac-7e6defe77733 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Trial without error: Towards safe reinforcement learning via human intervention
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258d8482-dc10-4668-9ef7-1a446d1fb73c · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ae7c4e-07b1-4d91-9455-667a6c5dbe44 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Smith, J
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47bd04e3-9676-486c-956f-e380a6bf15f0 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to be Safe: Deep RL with a Safety Critic
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbf0afc-f137-487a-8e8e-33a50aeca87e · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sim-to-real: Learning agile locomotion for quadruped robots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a91ff146-8d42-498a-bb9e-1893a6f293b2 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Reward Constrained Policy Optimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cdc0bd-9fc4-48a2-a58c-b8b6cc9060b6 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Emanuel Todorov, Tom Erez, and Yuval Tassa
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06645f40-c119-4ad3-a0dd-57568a859669 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mark Towers, Jordan K
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc01a51-20ed-4a4c-bfd2-176bcd8158ed · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2024/913
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6355c3b-f2b4-4c47-90bf-b0fef8aed300 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Kevin Zakka, Yuval Tassa, and MuJoCo Menagerie Contributors
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da7ee70-3e09-4f31-b7a0-41fbf86662a1 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2023/763
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2fe6394f-3268-4bcf-9cc2-11947860fb17 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Clipping bounds variance downstream of the singularity rather than removing it
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30d9b33d-1e35-4927-a0c1-db54f7ecacd9 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7879240e-bc99-45df-b474-a43106d4cbba · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281778e4-fe75-4cbe-b09f-56f2c4bb271f · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9722ad14-7b7e-44dd-89cf-5a72673eead8 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions CPO and PPO-Lagrangian additionally carry the constraint hyperparameters their objective requires (Section C.3)
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36000acf-9365-4be8-a4b0-19692dbfcb5d · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5209fcc6-7cba-41d5-8ac8-b87faf89dc3e · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2c3ada-c1ea-42b8-a5d1-ffe5a9adc628 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions global”). An alternative “per-segment
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ce74f8-3002-4757-8960-f03e60a71d89 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a347c806-4456-4019-bf06-729146d76549 · outbound
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.