Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:44:15.289881Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.01329.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:44:15.289881Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1eba9a00-639c-4083-85cb-f978864228db · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Towards Characterizing Divergence in Deep Q-Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd3b8ea-ca1a-401b-b5c1-521e33bd46b9 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning at the edge of the statistical precipice
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9012bf15-39fb-4c85-9e88-ae2930c1ae75 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Atari-5: Distilling the Arcade Learning Environment down to Five Games
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cb8308-9a86-4ddb-991d-a34a13b634e2 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Never give up: Learning directed exploration strategies, 2020
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c158a74f-ce46-4a51-ac71-a3108eda206b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unifying count-based exploration and intrinsic motivation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4b604a0-5918-4310-bf80-63c9e5ed83a8 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 714980d0-1038-4f32-8146-408f74c02da7 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Will Dabney, and R \' e mi Munos
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91486c0e-7de8-4ad9-9868-1988227725aa · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The theory of dynamic programming
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4caee5a4-d6dc-4201-a34c-10df449dec34 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Interference and generalization in temporal difference learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab5064c3-2b99-465c-9b1d-573ff2ac0daf · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation, 2018 a
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d30659ea-1eab-48c2-b419-4031954f0126 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f07d446-0420-4b67-8e80-e72e1a969202 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03652976-be24-49f7-bf0e-8ecf3ab81883 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b52edd0c-0814-40e9-be24-7cde3cdbe10e · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Phasic policy gradient
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88de9485-13b4-4789-8f86-17fa6b062881 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Loss of plasticity in deep continual learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a1084c3-0b78-49f7-b3a2-86d547419461 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Stop regressing: Training value functions via classification for scalable deep RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ec7e74d-589f-4bf7-8881-693668b66d7b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fujimoto, H
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40ed1020-c78a-49f5-b408-d40e0818c4a5 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a9231d9-a5f3-4db5-9681-80aabfe18382 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving performance in reinforcement learning by breaking generalization in neural networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 398644d4-ff48-4007-b3e2-65ca964f6c89 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Haarnoja, A
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c9e5f52-3bf3-4823-b7ec-01384df6b968 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning for model predictive control
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db07826b-be36-4c5a-baa9-4954c7e2210d · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Rainbow: Combining improvements in deep reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db39975b-1a29-41ba-870f-2a6b4559fffe · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Approximately optimal approximate reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e9f4d0f-9c87-423d-80bf-f327d0e7ff56 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Discor: Corrective feedback in reinforcement learning via distribution correction
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86f10e3c-19d3-43be-831d-836dc2b3b1bb · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9d82107-e6cb-4913-ae78-48ec50239c3b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Maxmin q-learning: Controlling the estimation bias of q-learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d0b4467-759a-436b-bcc0-6d29e243fd3c · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Optidice: Offline policy optimization via stationary distribution correction estimation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9e2609a-7fab-4c54-b406-4663a7ae6527 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02bf594d-fa09-4e4c-89fa-1bd7856647a7 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 692c1348-e985-40f7-b5ff-a4ec1e411a5d · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Understanding plasticity in neural networks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 296b198c-6a76-4466-a25c-824730fbc677 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Normalization and effective learning rates in reinforcement learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2f20bc-4c50-40a8-bf90-2a42869a932b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Human-level control through deep reinforcement learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a11509c-8cd1-46bb-8bef-8b133a160193 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? AlgaeDICE: Policy Gradient from Arbitrary Experience
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d31f03-dd55-440d-a529-5d9f1229d620 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c79c66-b119-417c-b4b2-d90458ca0b5e · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Nikishin, M
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 49e4a198-8e23-4210-a2a6-6a14c665cb30 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Mixtures of Experts Unlock Parameter Scaling for Deep RL
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation befae0ed-fecd-4bb2-92b0-127d5f9bb8d5 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Chatgpt: Optimizing language models for dialogue, 2022
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83f468a9-7d71-4521-bab2-62eb0ccd4f4b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Dota 2 with large scale deep reinforcement learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c23f6f7-da7b-4a29-bff5-dd379dde39c9 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Aaron van den Oord, and Remi Munos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 042f1dbe-316a-4063-aabf-d0de3c46f816 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The difficulty of passive learning in deep reinforcement learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0ae13c7-9db3-427e-84c4-660e1cd560dd · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Jha, Toshisada Mariyama, and Daniel Nikovski
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34226b10-1c78-4dad-a4a4-e40eae71538b · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fuzzy tiling activations: A simple approach to learning sparse representations online
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb97cefc-621b-43d2-a46e-78438b04dc6c · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Efros, and Trevor Darrell
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f8cf745-15fe-495c-8866-5c74e443db44 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bridging the gap between target networks and functional regularization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6832d5f-b75a-401a-b8b8-187ff913c706 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Decoupling value and policy for generalization in reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7edd74f-59f8-414a-98b4-47d23bbb7de1 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Prioritized experience replay
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c29176a-8c64-4fbb-8bf2-d94bd3197d30 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Schulman, S
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation baeb9d29-7246-4a99-ba62-f1f95bb9e9ef · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Proximal Policy Optimization Algorithms
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339911d0-db30-44f0-a4ee-5bfce67c664a · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Courville, Marc G
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0fce935c-1f8c-41d6-b479-817a1160e7f8 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f803f178-40de-4eb6-b7de-bdb65e7ad3b7 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? \#exploration: A study of count-based exploration for deep reinforcement learning, 2017
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5bdd617-acf5-4402-9c79-250644d4e6ae · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving deep reinforcement learning by reducing the chain effect of value and policy churn
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b92963fa-2d7e-45e5-8748-23f0767b9695 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning and td-gammon
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39dc3e1e-6ffe-45ae-b113-8b0222107181 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning with double q-learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d54ad17-5829-4f9c-a0f1-9b53f8d0d437 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep Reinforcement Learning and the Deadly Triad
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4711cce7-329c-4425-8ae4-aac846a25fca · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Overcoming the spectral bias of neural value approximation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac19e42c-0955-4e59-9db1-4a4516f45894 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60d6ec9f-97b9-48af-8fe1-a0a208a49cc9 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Learning invariant representations for reinforcement learning without reconstruction
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fdbd5858-aa05-4113-a705-2a16f45e2761 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Gendice: Generalized offline estimation of stationary values
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 719dedf7-6b7d-4556-ad52-758912305210 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Breaking the deadly triad with a target network
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76405fa5-0000-4990-93c4-51b2eff046fd · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? BeBold: Exploration Beyond the Boundary of Explored Regions
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4bfed1c-e7c4-4ae2-8c29-abd91c3ee521 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Noveld: A simple yet effective exploration criterion
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ac185bc-b5ee-42b2-a8b6-661b1dc03eea · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? @esa (Ref
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de190b9-18af-4c9c-a78a-2fde7f7fe7d6 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b258ad46-e590-4175-8533-77a1ac54c371 · outbound
Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.