Pith. sign in

Paper Citation Record · LEDGER

A Survey on Explainable Deep Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 7 inbound Pith citation observations for arXiv:2502.06869.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06869 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:22.041681Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:08:45.596041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:42.446579Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e81ce1a-1e00-4754-b65a-55f637c05cf9 · outbound

This paper cites CDT: Cascading Decision Trees for Explainable Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning CDT: Cascading Decision Trees for Explainable Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.993559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.993559Z digest=sha256:89e80c6e960ece118f920acb8ae822b2d318c03adc6ea7a5ea33c2d9f29bb018

Observation ee3e32a0-d58f-45bc-a7e4-2ae32c741b4f · outbound

This paper cites Empirical influence functions to understand the logic of fine-tuning.

A Survey on Explainable Deep Reinforcement Learning Empirical influence functions to understand the logic of fine-tuning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:22.150370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T19:17:22.016630Z digest=sha256:dc43a9891f1dea229d4da9ffe7b7bd60d8d9820e1cdd30f6f8a66ef60937f6bb

Observation ec979f0d-0b29-46bb-9a79-ab16384dd96d · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

A Survey on Explainable Deep Reinforcement Learning Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.027145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.027145Z digest=sha256:c4e2b6552020c65cc463963c9330a713fc27c46fadcca8f07984878a00a9fc01

Observation ad3a76e1-43dc-43d6-939d-1872bc260bf3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Survey on Explainable Deep Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.032072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.032072Z digest=sha256:401fb21a5218a891535cf7e86831e1175caaa7a61dea605e7127be5a60ef681b

Observation db1f2503-fc15-4377-983f-8d34455fe331 · outbound

This paper cites StarCraft II: A New Challenge for Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning StarCraft II: A New Challenge for Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.036708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.036708Z digest=sha256:008cf29ec38b1a98cf8dbbda786c77b9d8473869122c75b378ff346f798844f3

Observation 20ad6279-a0ce-4010-b33c-c23b7aeb18d5 · outbound

This paper cites Graying the black box: Understanding dqns.

A Survey on Explainable Deep Reinforcement Learning Graying the black box: Understanding dqns

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:22.261860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T19:17:22.041681Z digest=sha256:0f6f478656756e7f48d801932f04e84323d3cd7adf31fb34a2a1e810329e85d8

Observation bf3fe872-86b8-4ab3-85ea-41bf7df95140 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.022282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.022282Z digest=sha256:fd3f4952c69fc50270b4c0fc349f09832b5e069acebd50f30e3e93427fe8ed4c

Observation b8418f0f-c047-4149-90ec-39cef4f31e85 · outbound

This paper cites Do Influence Functions Work on Large Language Models?.

A Survey on Explainable Deep Reinforcement Learning Do Influence Functions Work on Large Language Models?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.010430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.010430Z digest=sha256:eb89db8512043384121a5e2152d44e8e6ea8b98dd798ba35095f32c018181f85

Observation a07730d5-54b2-44b5-9f9f-6985f2483595 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

A Survey on Explainable Deep Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.982058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.982058Z digest=sha256:f34e8e7d90ad511bed4d7406cc4991fbe21355bfdeefafaa67bb593f7aef1b1d

Observation 639bc1e6-02bb-4c76-aff6-82d189ed02b8 · outbound

This paper cites Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards.

A Survey on Explainable Deep Reinforcement Learning Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.005131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.005131Z digest=sha256:9abfb9917461db5135418e65c17878827d859ee81ca4dcc5216b205048339dc9

Observation 8e5c1f63-3ef3-46c3-a3c9-2ffc12ecd55e · outbound

This paper cites PromptExp: Multi-granularity Prompt Explanation of Large Language Models.

A Survey on Explainable Deep Reinforcement Learning PromptExp: Multi-granularity Prompt Explanation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.999381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.999381Z digest=sha256:d5bc4e27d8215b8a274c1c9e9e47516f6cf50f2d29fe7d5e01ee2af3729724b5

Observation c662d386-ef9e-417a-a2b9-e734c20dced9 · outbound

This paper cites Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models.

A Survey on Explainable Deep Reinforcement Learning Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.988056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.988056Z digest=sha256:d9dd6c44e219a1e78e82fafd0e61f67d513a92094ea955f4dedcabd9bbcb19a0

Pith citing papers

Observation 5feac328-cfef-44cd-9a0d-7f55b669f77a · inbound

Code-Driven Planning in Grid Worlds with Large Language Models cites this paper.

Code-Driven Planning in Grid Worlds with Large Language Models A Survey on Explainable Deep Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:45.596041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:45.596041Z digest=sha256:2ece084f7bf5655c7c241e59ea1fb02000b7fb21fb80e1d7bf1679b798a23182

Observation c0ba9a2d-127b-4f21-8246-65e015834008 · inbound

Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration cites this paper.

Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration A Survey on Explainable Deep Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:28.048575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:28.048575Z digest=sha256:4e0440eba0bc8536e64996348b368db5ad658b276a37751341beeabff8143325

Observation f13abb6d-7170-4bb9-be00-4edc0e906c68 · inbound

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI cites this paper.

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI A Survey on Explainable Deep Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:08:49.409444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:08:49.409444Z digest=sha256:15d387e84328c730e00bdf30f1d842f22a1fcec6f3c7938a5bd46272ee35a714

Observation 0712794e-c8e0-494b-bdb3-83cfa4fc9922 · inbound

Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments cites this paper.

Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments A Survey on Explainable Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:46:26.651930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:46:26.651930Z digest=sha256:294c4735dde9b15206a4e837610fa274e189a752d45a0f735b3c8eb87b92c854

Observation 59084918-d9da-4c68-8672-3270b04e54bd · inbound

M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs cites this paper.

M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs A Survey on Explainable Deep Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:20.792149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:42:03.298417Z digest=sha256:cfbd149a815e14a784dae03331f0f7d3ee1cfaf082281da1f6c5b4442ab8de69

Observation 537885f0-3e60-47b4-814f-40b0895d394c · inbound

Price of Fairness in Short-Term and Long-Term Algorithmic Selections cites this paper.

Price of Fairness in Short-Term and Long-Term Algorithmic Selections A Survey on Explainable Deep Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:11.156182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T10:04:04.849162Z digest=sha256:342995b4554defbb9b6669f2903769dc03184efd7619d58bcfb4d7ca5d677518

Observation 05e54b9f-39a5-4713-b01b-17b1f07fc09e · inbound

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI cites this paper.

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI A Survey on Explainable Deep Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:42.448163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T11:07:50.847698Z digest=sha256:b2e9c2dd04bd71fa87c7d754d29a17d48b13a88253d8add9be93c1658214d0b6