Pith. sign in

Paper Citation Record · LEDGER

Simplifying Deep Temporal Difference Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2407.04811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.04811 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:16.184323Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1cbf3a39-31ac-4f64-b5ea-4588a997c971 · inbound

Plasticity Loss in Deep Reinforcement Learning: A Survey cites this paper.

Plasticity Loss in Deep Reinforcement Learning: A Survey Simplifying Deep Temporal Difference Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T18:03:18.157374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T18:02:30.199552Z digest=sha256:bb3019e429f2f424aff031c2108521a75fcf0c8d562c02cbe82f7593b7c69cf4

Observation f23f055e-cb93-49fb-bd93-46373db78d0a · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:16.184323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:16.184323Z digest=sha256:95a832f1dde68c1195eb3b5ef144a630531d5c79bda166e384289228b8670a92

Observation f9805582-53ad-478a-984d-dfa78092c0f0 · inbound

Universal Value-Function Uncertainties cites this paper.

Universal Value-Function Uncertainties Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:55.363030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:49:55.363030Z digest=sha256:12dd1f09fc8cf5a00d55bd151019b07396419a832a0630bfe017ef02a4c9fad1

Observation 3349cc25-8617-4799-b1a8-445fb5b43491 · inbound

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control cites this paper.

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control Simplifying Deep Temporal Difference Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:33.176135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:33.176135Z digest=sha256:6a0d2c39c93ab0993f6bbe6d0298b1c97eef5151e75bde76a126b4cb4247b5c3

Observation f5e31804-8085-4696-9cd2-b998016c275e · inbound

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks cites this paper.

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:52.373117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:10:52.373117Z digest=sha256:011e68c34989c943e20793a4b83dcd79bf6a710a36a7ccbb3a4292505fec52c1

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · inbound

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models cites this paper.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:bf58493c7951f4b47bdd8c9c34516795fec53c05031a1ac147cfe511ebf5252e

Observation ab8b6e32-c3dd-4343-aa03-25fb4bda2d62 · inbound

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks cites this paper.

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:23.236918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:03:23.236918Z digest=sha256:c7c5b1be52e56b7f7554a5cb4632981e300d13ff039f162838bcc700fcab1e41

Observation 0a5445b3-83e8-45f8-bdab-b98429e1a043 · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Simplifying Deep Temporal Difference Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.239780Z digest=sha256:48d8ea84bdd10c99171434ed99860f1f18a241590ac4f0c32ad2d0736bb1e183

Observation 4c61f33e-1cc3-470c-b869-fd968e9c35dc · inbound

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning cites this paper.

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:57.233093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:57.233093Z digest=sha256:de14fa381cf647b5152b8164dc90ad2c59a60ecc4916250ead03308973f6f07c

Observation 1877c9b2-7022-41d7-9ccd-9958b2bb4981 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.774894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:2e68c3eff93ecd032db3d0e954ced9b21fa164fc2d0a7b19b4fa84d2f361a56b

Observation 18f36599-0f5f-4134-bf72-988583957402 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.429736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.429736Z digest=sha256:c8d74bc17e60d467d160bdc261f72b57fef90fd2f92d25fd6f2274356f4141bf

Observation 700bd6c9-d87a-4d83-b166-9073f9a2d975 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:49.847929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:7d7663d39b82396163201ded9e7ae9154ea9e94f3ec055aa7074c1c92c51fed8

Observation 66d0fc71-81f6-4aee-b9ff-a19f156460cf · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:12:41.290942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:e59f1a1702eea131b125e33dfc4ed65c11af1f04734dc410dd26023726253f74

Observation 99db9eed-557d-4a49-9d21-11518d8db892 · inbound

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations cites this paper.

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations Simplifying Deep Temporal Difference Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:36:26.179257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T10:27:55.337253Z digest=sha256:d5eafd77f027964fd8a0f7abfdc1aa09ac31e4d7d499be36c2c9b2a043efc14f

Observation a98f02b8-939f-483e-a53d-b0622f339fc3 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:24:27.808540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:50c7f4fe6bca5bf9338b22e24a8935d11e2626e358a24d6579cd174a92d79947

Observation 7fb16431-0e58-43f7-b718-674279663c0f · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.638679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:ade113d802ae7eb6b75dfd8b98f77636395bf6a12192aa567f843ca7d401dc3a

Observation e7d2bc9c-9f63-4cf8-a1f9-e1d79c416930 · inbound

Goal-Conditioned Agents that Learn Everything All at Once cites this paper.

Goal-Conditioned Agents that Learn Everything All at Once Simplifying Deep Temporal Difference Learning

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T05:00:21.532243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:59:48.867927Z digest=sha256:cc8c807f8652d64c0c1ef32b611c6cf6858820a0d2dcb8380b736a299c758cf5

Observation 610deba9-2a2d-4ea0-baf2-60ece410dbf5 · inbound

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion cites this paper.

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion Simplifying Deep Temporal Difference Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:05:49.709954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:54:12.099045Z digest=sha256:3b92b2f6112c0b4042f5a08f2a5d8978ab5af8964adf5b2dd50c6c0917bb6336

Observation 55d89c66-9695-4d48-b86b-6f1e692f7511 · inbound

Task diversity produces systematic transfer but inhibits continual reinforcement learning cites this paper.

Task diversity produces systematic transfer but inhibits continual reinforcement learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.240548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T18:58:56.444461Z digest=sha256:4c7e2041e04d979244ce44ed653e542431ed5a7fa37ca6e1881eddb18e6e7cc0

Observation 8fbaf87d-acd4-4039-b25b-b2f5d992d2c6 · inbound

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning cites this paper.

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:41:45.477737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:37:08.423086Z digest=sha256:eaec4feb591e71bc48d3eb931cfcb4d0aeab6a4fd6d7d48a9e748d5446766c29

Observation 03362496-b669-41a4-b8d4-ec3f5ad205a7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Simplifying Deep Temporal Difference Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.706984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:ea8ebb2c113a90a130dd5e2533cdd578c04dee1c4b229f1200d5d65c175c6c1e

Observation 3aef960c-934a-4938-8528-e304a4dd1541 · inbound

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning cites this paper.

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:51:17.007544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:51:17.007544Z digest=sha256:97fe10cec34720b301e7b9079a68d4a19d4b0a7f2f0d264c91145fa5978b6582