Pith. sign in

Paper Citation Record · LEDGER

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.12366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12366 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:34.992032Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff819525-73b8-41fa-b0f1-209cf67f3208 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.850630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.093584Z digest=sha256:22356827879cb7d6f98c4104248e04c0a8004f4f988324fac81cefc14e9c48e0

Observation 62341f45-5ae2-4e76-97d0-b1477f3800f1 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.154377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.154377Z digest=sha256:b994d4f8c154af07d7573f1a5665402cbadbd77c97094705e9db5c5d95820d01

Observation d9c691f7-9c49-429c-aa7c-8f33f9ce667d · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.717508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.238890Z digest=sha256:6ce1c791981257cfb93880fb53be7fd53510a10dd67d03737b259ebdd652bf8c

Observation bc84f432-7f24-4c6c-bdac-2375dc1ef1b7 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.317628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.317628Z digest=sha256:f6990724ab4b523a0f192e1ef557bdde4b987db7c14ae5f659be54a6e8e5297c

Observation 62b23a52-8c10-4b28-8f78-7cf8a2647170 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Solving Rubik's Cube with a Robot Hand

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.499519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.499519Z digest=sha256:40325085e65f026a90d2292b29c0b90ddd7bcfd2f8db9f3d35512ade76c40f5b

Observation 3628204b-3c6f-4f7e-aed3-8e1bc22dc60a · outbound

This paper cites Assessing the Impact of Distribution Shift on Reinforcement Learning Performance.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Assessing the Impact of Distribution Shift on Reinforcement Learning Performance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.540733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.540733Z digest=sha256:ed3931a18cff56281563d99ab7f2fae97895b8fd111400c74376bcb94e639d78

Observation d8067b51-5833-42ea-bdd8-8ddf3eaa0c66 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.560428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.636171Z digest=sha256:aec0e497445715806e9ac0a6df84f1cb8b5f9601b4cf8e6ff44d7245ae1170ca

Observation 2f0afb74-432c-4faa-af13-36455bc8470c · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.681317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.681317Z digest=sha256:9fa72e5375bf9ecf5c0b494e760af771c549f5cb11008752c48c0d200c3c688f

Observation eae48b2e-276e-42bd-a27d-a3ad041d95a6 · outbound

This paper cites & Krueger, D.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Krueger, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:37.422518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.765659Z digest=sha256:d547e41333c02f146e3258b457bfcaa0abae36d3dcd9af1dca71f861e85fa8d6

Observation b85e1bf7-f3de-45f9-ab9b-376693f31989 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.229349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.800807Z digest=sha256:e7e0be18439b28ab5c65152e3e4cc6a59ebb612270de99800fdc1b16f21ecafd

Observation f4514670-e34f-4f9d-8d03-cc0053e9feb2 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.062856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.812768Z digest=sha256:484696d63a5b2b163419f065a24ab4d6dbb6babb4fd899ee19ce6cb4f74a27ab

Observation 1b0bfa77-eac1-43fc-8dea-8a54105fb0c4 · outbound

This paper cites Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:35.277560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.844964Z digest=sha256:2f3b07bc3f5123b2df01ae30716f2ab04a09df30eaa5f2fb141deae80ed0a045

Observation cae01744-3e5b-4045-b837-214e3d193530 · outbound

This paper cites & Fal- cone, F.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Fal- cone, F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.900260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:33.959059Z digest=sha256:3c01e7b13cf9ee363f88708de4c98fac45efdd5043801c2f95ae7e6675a6d1b7

Observation 4a8df109-521d-4be5-b39c-fc5d20ccc265 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:36.705638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.070052Z digest=sha256:1936f21f36ccab0a327d1ebe3861ebd02aae61b1f1c82f49b3924124820ef2d9

Observation 05b5c852-5616-4152-b1c2-c744954b8c4e · outbound

This paper cites ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024).

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.547757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.202429Z digest=sha256:7ff37aff92874362bb003bd14ff5fecf7ba957cac09375bebcc8222bac7805a4

Observation 24ff9008-bd64-4e78-9ffa-eaa4f21ff008 · outbound

This paper cites Simulated augmented reality and virtual immersive technologies.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Simulated augmented reality and virtual immersive technologies

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.419593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.334632Z digest=sha256:c12ed51aaf4be4cb092a8397de1acc6415d979ec8fdfa77bb907ff773aab5dc0

Observation 2acbf66a-3a1c-4421-8c16-87365c2f07d0 · outbound

This paper cites & Yablon, Z.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Yablon, Z

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.257621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.464284Z digest=sha256:72cf9cf66f3e68a9f96cf1ccc5dcb7da463db1295f7b9017530cbef91ebf8bbf

Observation 60cc0211-e0a1-4795-80a4-b7c81e549c6c · outbound

This paper cites & Luo, X.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Luo, X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.097045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.633125Z digest=sha256:a1ae59d2786c10e9c2c92bdb39d7bfc0302e6682fbc17d8be1c3336c9acdbf12

Observation c28be68d-25d9-4e01-bd62-8f058b35881b · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:35.885780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.746730Z digest=sha256:f0c78d28fc9ef735a34993ff7bdd7851f118bbf132648d2af7a3cad3f2b29bfd

Observation cb1e6732-4b91-4ebf-a565-2c1f2eb8f6a3 · outbound

This paper cites & Lee, J.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Lee, J

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:35.649468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:54:34.847455Z digest=sha256:cffa2f73d8e0e3341b1303e273108b75b8582f636897a1f53ca0291f8e29f28c

Observation 6735d5a7-b288-4389-a2a1-463332410c9d · outbound

This paper cites Unity: A General Platform for Intelligent Agents.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unity: A General Platform for Intelligent Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:34.992032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:34.992032Z digest=sha256:c200c2aa2ecee0915223daff7bf56dced84aa8db70e15150841282a3eb63f8c3

Pith citing papers

No inbound Pith citation observations are available.