Pith. sign in

Paper Citation Record · LEDGER

Is Exploration or Optimization the Problem for Deep Reinforcement Learning?

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.01329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01329 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:44:15.289881Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1eba9a00-639c-4083-85cb-f978864228db · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Towards Characterizing Divergence in Deep Q-Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.402366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.402366Z digest=sha256:ca5a6adaf5efcac3decc9072cfe6cf71985ace340e5d6fc33b2d6e5ad0230c86

Observation 0bd3b8ea-ca1a-401b-b5c1-521e33bd46b9 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning at the edge of the statistical precipice

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.445434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.523139Z digest=sha256:4cd682d4a56db280a53990c52653424949c391f9832130a7dd77641e83f5e401

Observation 9012bf15-39fb-4c85-9e88-ae2930c1ae75 · outbound

This paper cites Atari-5: Distilling the Arcade Learning Environment down to Five Games.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Atari-5: Distilling the Arcade Learning Environment down to Five Games

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.701517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.701517Z digest=sha256:29377488c008cf974b249b37f60337804dd90b4671043ec4c5f03e151d6de0d7

Observation 65cb8308-9a86-4ddb-991d-a34a13b634e2 · outbound

This paper cites Never give up: Learning directed exploration strategies, 2020.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Never give up: Learning directed exploration strategies, 2020

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.429262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.835102Z digest=sha256:87d021510de5da7fd4bca8aeb445eaaeca1da9452c0b00c01767d7c2ece55b8f

Observation c158a74f-ce46-4a51-ac71-a3108eda206b · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unifying count-based exploration and intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.414088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.997083Z digest=sha256:70cfe554b6b627a51e518713aac689055406e20e97b831e96b80a2413d8fd0bc

Observation e4b604a0-5918-4310-bf80-63c9e5ed83a8 · outbound

This paper cites Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T05:44:15.608184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.117380Z digest=sha256:7fa03e3450c0217150e634ed9b99952fed65b4e2e162770c7fc9b3ed7cb64a39

Observation 714980d0-1038-4f32-8146-408f74c02da7 · outbound

This paper cites Bellemare, Will Dabney, and R \' e mi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Will Dabney, and R \' e mi Munos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.399301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.284037Z digest=sha256:169c7e71dba94e9361f795fdf2422913ca6b0f317258efcd8da25558d299cbb4

Observation 91486c0e-7de8-4ad9-9868-1988227725aa · outbound

This paper cites The theory of dynamic programming.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The theory of dynamic programming

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.384753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.418889Z digest=sha256:d48bc0891d1f52948811857ab3b61434ea804240d4dec42b8778b8172f691bde

Observation 4caee5a4-d6dc-4201-a34c-10df449dec34 · outbound

This paper cites Interference and generalization in temporal difference learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Interference and generalization in temporal difference learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.369073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.518429Z digest=sha256:57b4713572a6c28fa3339c1857453fdb4d4710d2e88ccb5bf04aac5d18973773

Observation ab5064c3-2b99-465c-9b1d-573ff2ac0daf · outbound

This paper cites Exploration by random network distillation, 2018 a.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation, 2018 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.353804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.624275Z digest=sha256:803c2fe0a08b6b90231a12990e6e2c27b875398800b1f43f919831a6c16e0037

Observation d30659ea-1eab-48c2-b419-4031954f0126 · outbound

This paper cites Exploration by random network distillation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.339315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.800901Z digest=sha256:41e5e2762fd0089163848844043860288cb5c99de9789a545948b0b788afa5af

Observation 8f07d446-0420-4b67-8e80-e72e1a969202 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.323706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.940884Z digest=sha256:a43c18db0a35bcdae94d6c74ec1ae28093aa0106f56ec98c53d1403ce1ddff65

Observation 03652976-be24-49f7-bf0e-8ecf3ab81883 · outbound

This paper cites Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.474846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.114051Z digest=sha256:df724517006146504740c6ba074952ea2bd31b6ef2335dc1bef9b5dcc1facd8d

Observation b52edd0c-0814-40e9-be24-7cde3cdbe10e · outbound

This paper cites Phasic policy gradient.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Phasic policy gradient

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.308603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.226971Z digest=sha256:7d0549162193031ce021cdf1a85829ac2c96146a93bc7fe035273c40fe9b2828

Observation 88de9485-13b4-4789-8f86-17fa6b062881 · outbound

This paper cites Loss of plasticity in deep continual learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Loss of plasticity in deep continual learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.291200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.362972Z digest=sha256:321433a265169bf3d75bf89666b4f457bb86f1643000545b1a7f6646efe3b94a

Observation 5a1084c3-0b78-49f7-b3a2-86d547419461 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Stop regressing: Training value functions via classification for scalable deep RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.273604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.464882Z digest=sha256:7bdbbe10a17fe29f5364cab10b8c9af68b72affdac40662e540ae9c1956f62db

Observation 5ec7e74d-589f-4bf7-8881-693668b66d7b · outbound

This paper cites Fujimoto, H.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fujimoto, H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.257618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.647048Z digest=sha256:13d901f262052de6bca67e62ef94b2964d2cc5ba91be332c42c037e72b440fbf

Observation 40ed1020-c78a-49f5-b408-d40e0818c4a5 · outbound

This paper cites Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.757638Z digest=sha256:7e65e5c36b486fb2ffdaba3f02c5b9de3c8414feb4a548af9009187c46417f08

Observation 0a9231d9-a5f3-4db5-9681-80aabfe18382 · outbound

This paper cites Improving performance in reinforcement learning by breaking generalization in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving performance in reinforcement learning by breaking generalization in neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.242508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.931406Z digest=sha256:f09ca0b58a418e1934bccddca1e77ff756ea7e859902ebcd784919a183cdaf9a

Observation 398644d4-ff48-4007-b3e2-65ca964f6c89 · outbound

This paper cites Haarnoja, A.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Haarnoja, A

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.224251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.108446Z digest=sha256:1c50acbbdceaa9bd6ef268f211384b103737b9d5f99f0ad1a0e6c079d54be223

Observation 3c9e5f52-3bf3-4823-b7ec-01384df6b968 · outbound

This paper cites Temporal difference learning for model predictive control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning for model predictive control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.208501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.243178Z digest=sha256:82720bc3f5749b3299fe7be04fcd556e77687d16ae9a39058f312a190b85703a

Observation db07826b-be36-4c5a-baa9-4954c7e2210d · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Rainbow: Combining improvements in deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.191803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.406629Z digest=sha256:98d9cefe16598704fbf4b04f06bdf27f6fcb3b00838920e80154f6d720f74e94

Observation db39975b-1a29-41ba-870f-2a6b4559fffe · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Approximately optimal approximate reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.174366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.537137Z digest=sha256:8a4e109ef7ec60a95994abc0218e05a904da024c5cf8c24901f2ed4cf4b07433

Observation 8e9f4d0f-9c87-423d-80bf-f327d0e7ff56 · outbound

This paper cites Discor: Corrective feedback in reinforcement learning via distribution correction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Discor: Corrective feedback in reinforcement learning via distribution correction

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.158893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.662962Z digest=sha256:693c84e4e150ba13911edf2a14e5a7269c6641a966bf0a907a4261a90158ab09

Observation 86f10e3c-19d3-43be-831d-836dc2b3b1bb · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.141186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.793980Z digest=sha256:b5b540b1632b9b31db22cfddf3f661c2acba4e8b90f6b9c4c2813419d71a2efd

Observation b9d82107-e6cb-4913-ae78-48ec50239c3b · outbound

This paper cites Maxmin q-learning: Controlling the estimation bias of q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Maxmin q-learning: Controlling the estimation bias of q-learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.125073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.974012Z digest=sha256:ebaae3a745da027cc2fff0b0185b938770af3cb620e3ba9f79da723e21230f90

Observation 0d0b4467-759a-436b-bcc0-6d29e243fd3c · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.108519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.071409Z digest=sha256:6477f8ddcb58cfd43cc93704f024497c29266c9a5857f10065fcdd960767fa7f

Observation a9e2609a-7fab-4c54-b406-4663a7ae6527 · outbound

This paper cites Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.092348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.090226Z digest=sha256:5599680502942edfb35cbbbd20ce0e5f0193036cf92a4558fa882f2199b0f323

Observation 02bf594d-fa09-4e4c-89fa-1bd7856647a7 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.076826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.102829Z digest=sha256:59c2a21985268ccfb272444d1ac6738e33c3b4399aec9c291a557a037d73b3d6

Observation 692c1348-e985-40f7-b5ff-a4ec1e411a5d · outbound

This paper cites Understanding plasticity in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Understanding plasticity in neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.060939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.107894Z digest=sha256:a37702c4726ffdd9653ab040e3d2d2011f35ee2abc6f1943811b4031243ec1f3

Observation 296b198c-6a76-4466-a25c-824730fbc677 · outbound

This paper cites Normalization and effective learning rates in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Normalization and effective learning rates in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.112426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.112426Z digest=sha256:aed91635fec742a94190aaa485e2dae4b1b8b612aed7472370246f15864318d0

Observation 5c2f20bc-4c50-40a8-bf90-2a42869a932b · outbound

This paper cites Human-level control through deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Human-level control through deep reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.117368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.117368Z digest=sha256:d3933c799d440675879a522685a1b3ab7483748ec1b1084cafcff1036ad6d063

Observation 8a11509c-8cd1-46bb-8bef-8b133a160193 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.122140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.122140Z digest=sha256:6c6e62f1f6bc311fb84239f6af47d5cee80c7025025f04f770beca704b256202

Observation 64d31f03-dd55-440d-a529-5d9f1229d620 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.127132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.127132Z digest=sha256:f72a3b4f88503f12e012830ad82e8e724c7de6726ef827874b5297671e263b04

Observation 73c79c66-b119-417c-b4b2-d90458ca0b5e · outbound

This paper cites Nikishin, M.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Nikishin, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.032454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.132767Z digest=sha256:9df98c36b30cabfc634d2eb0429444b68ac7ed4577e8f92741156d12f2b7ffbd

Observation 49e4a198-8e23-4210-a2a6-6a14c665cb30 · outbound

This paper cites Mixtures of Experts Unlock Parameter Scaling for Deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Mixtures of Experts Unlock Parameter Scaling for Deep RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.137915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.137915Z digest=sha256:2170b14a5508880040ac98628183e8bfb430cc1bff039d3289d0c6d177a42eaa

Observation befae0ed-fecd-4bb2-92b0-127d5f9bb8d5 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2022.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Chatgpt: Optimizing language models for dialogue, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.016932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.142696Z digest=sha256:edba7d78342a8727d63ee3c8ae2d78841a010acb5835a52b578d45c746f0a4db

Observation 83f468a9-7d71-4521-bab2-62eb0ccd4f4b · outbound

This paper cites Dota 2 with large scale deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Dota 2 with large scale deep reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.002008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.147758Z digest=sha256:56c647c287eb31b0af9b209a3af19d9ce23e5ea3bd70e4a0eff9b2c314bb32a1

Observation 6c23f6f7-da7b-4a29-bff5-dd379dde39c9 · outbound

This paper cites Bellemare, Aaron van den Oord, and Remi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Aaron van den Oord, and Remi Munos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.987483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.152329Z digest=sha256:2cc5f2c174e35a19dff9df3d4a2fd27ab4704b14300389980636b1f7110b2b27

Observation 042f1dbe-316a-4063-aabf-d0de3c46f816 · outbound

This paper cites The difficulty of passive learning in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The difficulty of passive learning in deep reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.972739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.157046Z digest=sha256:1a0c29e0c90993815f6f86cfccbd8b32421dc575c7c28749f01de6b78d0fcf0e

Observation d0ae13c7-9db3-427e-84c4-660e1cd560dd · outbound

This paper cites Jha, Toshisada Mariyama, and Daniel Nikovski.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Jha, Toshisada Mariyama, and Daniel Nikovski

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.956169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.161903Z digest=sha256:1ea37eac84b75ad35cf1d30a9bc99d931bce57bd8aa9ff11a61323eddd503475

Observation 34226b10-1c78-4dad-a4a4-e40eae71538b · outbound

This paper cites Fuzzy tiling activations: A simple approach to learning sparse representations online.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fuzzy tiling activations: A simple approach to learning sparse representations online

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.940986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.167501Z digest=sha256:79abfc9a704a51320566f7a51b9c0107e6d8f88e5410505575d6c395af110b3f

Observation eb97cefc-621b-43d2-a46e-78438b04dc6c · outbound

This paper cites Efros, and Trevor Darrell.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Efros, and Trevor Darrell

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.926529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.172719Z digest=sha256:57f91a000aba96ff014599b0f582f17c897e0f52c1b17dfff88bdd96e7d50e3a

Observation 3f8cf745-15fe-495c-8866-5c74e443db44 · outbound

This paper cites Bridging the gap between target networks and functional regularization.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bridging the gap between target networks and functional regularization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.909906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.179837Z digest=sha256:53c1749ddd22abf3c9dc1ff7f8a135f72bbecc721a904a686756fbbe02e3286a

Observation d6832d5f-b75a-401a-b8b8-187ff913c706 · outbound

This paper cites Decoupling value and policy for generalization in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Decoupling value and policy for generalization in reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.895196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.185252Z digest=sha256:37f86d7d54ed14cec611342fb48f8fa6dd6caf9a5c91ea82d2572d78df366ebe

Observation b7edd74f-59f8-414a-98b4-47d23bbb7de1 · outbound

This paper cites Prioritized experience replay.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Prioritized experience replay

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.880540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.190676Z digest=sha256:3785b08ac2437b440b9724f62ff66528f4072ae1ed30178288dbc0a3d89108da

Observation 9c29176a-8c64-4fbb-8bf2-d94bd3197d30 · outbound

This paper cites Schulman, S.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Schulman, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.864276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.196969Z digest=sha256:f2f667a701f3ffbf3d638897067eeebb5f4eded9bc6e8cfefbc725aa96cb86b2

Observation baeb9d29-7246-4a99-ba62-f1f95bb9e9ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Proximal Policy Optimization Algorithms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.201847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.201847Z digest=sha256:5078bc38d68192dee02a4c88b8cb885a925ead0685882884663cc337839ea63a

Observation 339911d0-db30-44f0-a4ee-5bfce67c664a · outbound

This paper cites Courville, Marc G.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Courville, Marc G

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.849518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.206930Z digest=sha256:8eacfabb0a16dac5abc4de2d25585eef9e70e4d5ef09d93b0d7827329d7e052d

Observation 0fce935c-1f8c-41d6-b479-817a1160e7f8 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.834305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.212085Z digest=sha256:dcd4d2b0334f646b81a44f0e8dceffd98c43e8f24a7f09a8289f25a6403562f9

Observation f803f178-40de-4eb6-b7de-bdb65e7ad3b7 · outbound

This paper cites \#exploration: A study of count-based exploration for deep reinforcement learning, 2017.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? \#exploration: A study of count-based exploration for deep reinforcement learning, 2017

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.817834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.216775Z digest=sha256:24d861454c6bd9d675c6c1f2e660bed6c0d7401b71c92d185e57fb1a0db4b47b

Observation b5bdd617-acf5-4402-9c79-250644d4e6ae · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.802566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.221980Z digest=sha256:965ca7d592ef8e734bc9d7b9eb60a99e86e0a6d7d0519bf89757647b1b817074

Observation b92963fa-2d7e-45e5-8748-23f0767b9695 · outbound

This paper cites Temporal difference learning and td-gammon.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning and td-gammon

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.227056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.227056Z digest=sha256:743df7b317bcd8cb6d04700651866750900c7a151de10d4f3e926123bb846e4b

Observation 39dc3e1e-6ffe-45ae-b113-8b0222107181 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning with double q-learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.232608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.232608Z digest=sha256:561379131c6cc9b0b74c486bcc1883e4d22f9f4539667e521375fd31f997f001

Observation 4d54ad17-5829-4f9c-a0f1-9b53f8d0d437 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep Reinforcement Learning and the Deadly Triad

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.237995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.237995Z digest=sha256:4c91f73b9fdcf002e8851a9662802ac14f87b79bcf6513529ee579b6b024bbdd

Observation 4711cce7-329c-4425-8ae4-aac846a25fca · outbound

This paper cites Overcoming the spectral bias of neural value approximation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Overcoming the spectral bias of neural value approximation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.766412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.242757Z digest=sha256:1ed619f33023814e2ddfb1a2a434ec04951a541541f0153a9cc503ae7a1992ad

Observation ac19e42c-0955-4e59-9db1-4a4516f45894 · outbound

This paper cites MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.749648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.247375Z digest=sha256:bfd2464bb65cc7c1635eebd69e33651bb41b11b8c97b17da43922569b8151644

Observation 60d6ec9f-97b9-48af-8fe1-a0a208a49cc9 · outbound

This paper cites Learning invariant representations for reinforcement learning without reconstruction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Learning invariant representations for reinforcement learning without reconstruction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.734634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.251844Z digest=sha256:bfbd92c887e84d0dcdad83f2930f25369006e6b22d5d939f5742b4017ce96527

Observation fdbd5858-aa05-4113-a705-2a16f45e2761 · outbound

This paper cites Gendice: Generalized offline estimation of stationary values.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Gendice: Generalized offline estimation of stationary values

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.719434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.256785Z digest=sha256:07b3075ff0c114363f67d26b2044501899eb950ab6f064b1322d30135294a928

Observation 719dedf7-6b7d-4556-ad52-758912305210 · outbound

This paper cites Breaking the deadly triad with a target network.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Breaking the deadly triad with a target network

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.704127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.262635Z digest=sha256:7bfefc0d889021cd1b6c6ec42b247401940edcc6e5efb44b5a3f802eaa0ed348

Observation 76405fa5-0000-4990-93c4-51b2eff046fd · outbound

This paper cites BeBold: Exploration Beyond the Boundary of Explored Regions.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? BeBold: Exploration Beyond the Boundary of Explored Regions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.267730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.267730Z digest=sha256:3faea359c46a8477ae6153a2c7618956cad3ec50975571d10a4d6d5d91192821

Observation d4bfed1c-e7c4-4ae2-8c29-abd91c3ee521 · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Noveld: A simple yet effective exploration criterion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.689132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.273651Z digest=sha256:cf4e2c36dd2a82c6e387d096d219d65048ae61351795a5384a02dc1a06a70699

Observation 3ac185bc-b5ee-42b2-a8b6-661b1dc03eea · outbound

This paper cites @esa (Ref.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.279905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.279905Z digest=sha256:9d5964621e3feac19e848a7793fd62f300c5c5d60c30b8cd3b68cea051a244f0

Observation 2de190b9-18af-4c9c-a78a-2fde7f7fe7d6 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.285200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.285200Z digest=sha256:5e07f8a14e3579cebb7c1cf411d8e0c8017fc5d81dec95600986b4b591cb7cd2

Observation b258ad46-e590-4175-8533-77a1ac54c371 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.653607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.289881Z digest=sha256:669e7337b9bf14c4b123e235fd8e3ea4af95e184ca1d31e1ad45dccb9983226c

Pith citing papers

No inbound Pith citation observations are available.