Pith. sign in

Paper Citation Record · LEDGER

Simplifying Deep Temporal Difference Learning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2407.04811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.04811 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:06:12.218071Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1cbf3a39-31ac-4f64-b5ea-4588a997c971 · inbound

Plasticity Loss in Deep Reinforcement Learning: A Survey cites this paper.

Plasticity Loss in Deep Reinforcement Learning: A Survey Simplifying Deep Temporal Difference Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T18:03:18.157374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T18:02:30.199552Z digest=sha256:d35b4b669587ed8fd1a96c1015b25427bdf315565516762230157aaf56d08526

Observation 62afaa51-dc9b-4aa3-a093-68423ef2eef3 · inbound

Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation cites this paper.

Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation Simplifying Deep Temporal Difference Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:18.111207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:18.111207Z digest=sha256:0a7a79dc1b7d0f5449ce896e18ad6322915c4fafc5b8b7dd5d67995869e461a0

Observation ca7628d2-bd2c-4038-abc2-85d805080b24 · inbound

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles cites this paper.

Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles Simplifying Deep Temporal Difference Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:06:12.218071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:06:12.218071Z digest=sha256:3a23c7d646ae680288d22fbd44658c249a26b18632bde56588b7877c276f5b88

Observation f23f055e-cb93-49fb-bd93-46373db78d0a · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:16.184323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:16.184323Z digest=sha256:b439edd1952cd4e4176559389185bb0a77fd3e31de77f779f2bb48da533e3e77

Observation f9805582-53ad-478a-984d-dfa78092c0f0 · inbound

Universal Value-Function Uncertainties cites this paper.

Universal Value-Function Uncertainties Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:55.363030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:49:55.363030Z digest=sha256:a6eb8c5ff636fe135c0541c767d043a7be50421f9073fe69c8e1e45b662ec313

Observation 3349cc25-8617-4799-b1a8-445fb5b43491 · inbound

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control cites this paper.

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control Simplifying Deep Temporal Difference Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:33.176135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:33.176135Z digest=sha256:c4663765678fcf712942d33feeabb96acd118ad258aae5d810dfd4288e62e7b5

Observation f5e31804-8085-4696-9cd2-b998016c275e · inbound

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks cites this paper.

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:52.373117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:10:52.373117Z digest=sha256:d921a625036c3cb1a5890df02a7e7c388bd922c2de71e825806318000b832908

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · inbound

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models cites this paper.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:64b54385634984661aa7c59ef559c3c339f16fc70d3c2cfac04a8eb2b89be45c

Observation ab8b6e32-c3dd-4343-aa03-25fb4bda2d62 · inbound

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks cites this paper.

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:23.236918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:03:23.236918Z digest=sha256:ce44034a9f3e750fd881dcee8e78aab54f7d08e1d4f5d5b7e0444782262efe07

Observation 0a5445b3-83e8-45f8-bdab-b98429e1a043 · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Simplifying Deep Temporal Difference Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.239780Z digest=sha256:2c9289df11a647d9e8aba0590f7cb2fd55dd494330ce241ae1ef6727d814ebc0

Observation 4c61f33e-1cc3-470c-b869-fd968e9c35dc · inbound

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning cites this paper.

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:57.233093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:57.233093Z digest=sha256:b2a7bcf8bf2d5b6b86ece79e62336104fc1e31c9f40ed376ecaadbaed20e461e

Observation 1877c9b2-7022-41d7-9ccd-9958b2bb4981 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.774894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:3f29bd1a1821ed194ab38df46aa240a2fc6da478d40eb30f3ba6dc94497b7c21

Observation 18f36599-0f5f-4134-bf72-988583957402 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.429736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.429736Z digest=sha256:bc87f1264a12643153302d8d6e7e3e87594e8634887b2dea84a63a78fcc9a4de

Observation 700bd6c9-d87a-4d83-b166-9073f9a2d975 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:49.847929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:94a99a81eee403c16b90f338ce733b07f0517e6893a725dc0e36c62855a9c707

Observation 66d0fc71-81f6-4aee-b9ff-a19f156460cf · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:12:41.290942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:d1f3aa1951737a4bdec33aa830d1f791c664a5e3b153d2b6e523c2a226947e1d

Observation 99db9eed-557d-4a49-9d21-11518d8db892 · inbound

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations cites this paper.

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations Simplifying Deep Temporal Difference Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:36:26.179257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T10:27:55.337253Z digest=sha256:60ca3698fdf8b4ec0a2f6fad5e9bae36e0cd3e094545ec61df0586487cb706b2

Observation a98f02b8-939f-483e-a53d-b0622f339fc3 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:24:27.808540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:86390eb46d0b6a997d4b3fdeb7b10b54cc2ab5ca4ed33408d02f79993967278a

Observation 7fb16431-0e58-43f7-b718-674279663c0f · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.638679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:df5c8d27bbe0e0a0229631ecffabda81fcdeb37860d3b340777f980b2be1b878

Observation e7d2bc9c-9f63-4cf8-a1f9-e1d79c416930 · inbound

Goal-Conditioned Agents that Learn Everything All at Once cites this paper.

Goal-Conditioned Agents that Learn Everything All at Once Simplifying Deep Temporal Difference Learning

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T05:00:21.532243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-25T04:59:48.867927Z digest=sha256:e3d26acf54c1e629a640ad952c6f205d95d73392c160769de9695f187576f957

Observation 610deba9-2a2d-4ea0-baf2-60ece410dbf5 · inbound

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion cites this paper.

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion Simplifying Deep Temporal Difference Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:05:49.709954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T00:54:12.099045Z digest=sha256:57e86866e3b98b1f70f03419d1ab690e969545b1ca8d88080cf0d0847eabf5c9

Observation 55d89c66-9695-4d48-b86b-6f1e692f7511 · inbound

Task diversity produces systematic transfer but inhibits continual reinforcement learning cites this paper.

Task diversity produces systematic transfer but inhibits continual reinforcement learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.240548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T18:58:56.444461Z digest=sha256:b0cb966c82432097d9a0678f31ed9bf46cd5fab747e1a99ba78fd911451c8512

Observation 8fbaf87d-acd4-4039-b25b-b2f5d992d2c6 · inbound

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning cites this paper.

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:41:45.477737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:37:08.423086Z digest=sha256:cc4cadc5195341b7468b768ae8e4a13fc1dd3c00930c54586b7f0e8f2a95824d

Observation 03362496-b669-41a4-b8d4-ec3f5ad205a7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Simplifying Deep Temporal Difference Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.706984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:40f40c0f9a0f610164ba306c986b739d20beab8e993cc2edfd1f2ecb8058def8

Observation 3aef960c-934a-4938-8528-e304a4dd1541 · inbound

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning cites this paper.

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:51:17.007544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:51:17.007544Z digest=sha256:87d1e320120d09ef03c6f53fbb76d697da830c3bde12249be329bf2726cdf322

Observation 77292a3b-3132-4cd7-9d29-a948cd53174e · inbound

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks cites this paper.

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks Simplifying Deep Temporal Difference Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T10:11:39.377230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T10:11:39.377230Z digest=sha256:0c33e0ee83dd128fd8c4bee8833994675124dbf6cda11ed89ee603524f2f6979

Observation 8229144f-b47a-414d-96d1-c6e1f821f108 · inbound

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks cites this paper.

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks Simplifying Deep Temporal Difference Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:45:49.213057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:45:49.213057Z digest=sha256:99cdd48e16e4fd06f99971ebdbeb93fcb2b34c952073c21ad804b7889fc1270d

Observation 3ee3e547-68dc-4b97-96fb-391fe6fc8959 · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Simplifying Deep Temporal Difference Learning

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:45.608194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:45.608194Z digest=sha256:411939895bf7e60ae13a87407a88b99b1d923e36f41b413d165910bab706340a