Pith. sign in

Paper Citation Record · LEDGER

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2606.21297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.21297 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T14:39:59.778081Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:38.157587Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 5f317c91-71b0-40a7-ab9f-9a970e2d1518 · outbound

This paper cites Sutton and Andrew G.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Sutton and Andrew G

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:e2b0f697b435ba8f0853a185e037006a38952efb93697e4ec3fbe6124bec260a

Observation a0d3fbe8-369a-4054-86f1-67e6a929692c · outbound

This paper cites Human-level control through deep reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Human-level control through deep reinforcement learning

Reference 2

Resolution
metadata mismatch
doi, observed 2026-06-26T14:49:32.407224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:26f55f1c89502e7a02c6852c0e65130c82f8bd61377f342769af212de5cdd39a

Observation d8689425-c486-4cf9-aca1-dd06a3717f56 · outbound

This paper cites Proximal Policy Optimization Algorithms.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Proximal Policy Optimization Algorithms

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.625838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b1c36d9fe599d2c20f61b169718b94165cec5a1d312c0729193a408040f02e51

Observation 7983f709-9b14-42f0-9b66-39c18e85c3f3 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Addressing function approximation error in actor-critic methods

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:c133e3dda7a3801db13f414ab4bb90ca4c5127afdc8f89767ce43a3d4dc247a0

Observation be50982b-608b-4d1c-bcb1-f7387dd02a94 · outbound

This paper cites URLhttp://proceedings.mlr.press/v80/fujimoto18a.html.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning URLhttp://proceedings.mlr.press/v80/fujimoto18a.html

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:3035f378baa4009ee9d71898c41e001ced59c5a6e5975d6f93cbad6393fffc05

Observation 2cedea32-7028-4c93-b2a7-caaee29a8014 · outbound

This paper cites Lillicrap, Jimmy Ba, and Mohammad Norouzi.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Lillicrap, Jimmy Ba, and Mohammad Norouzi

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:313c1a00ef0a766d9c1581ee0e788331c04aa27a620995b89e25d02ce925fd92

Observation 2ac79882-3c57-4cfa-81d1-09b962c16e4f · outbound

This paper cites Mastering Diverse Domains through World Models.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Mastering Diverse Domains through World Models

Reference 7

Resolution
malformed identifier
local_arxiv, observed 2026-06-26T14:49:32.409840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:f7cf12ad5e275965cf69645e88a775d93fbcdd94e2eb980a0c8b8bd4ec3f5341

Observation 9e1c918e-9b3f-4449-8019-7c8eb6290035 · outbound

This paper cites Temporal difference learning for model predictive control.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Temporal difference learning for model predictive control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:456a8b177db41556c95aabf837bb80e2aa9827d62e453c93adad9e4bbd862731

Observation cab8bb7a-a644-41ef-b3fe-61a3a37a1fc0 · outbound

This paper cites TD-MPC2: scalable, robust world models for continuous control.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning TD-MPC2: scalable, robust world models for continuous control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b5f7d8081b308bfa331f4d6bd1ec895cbfd0071e1b40cd233c3d42e747a11ce3

Observation d762848a-12b6-45b7-b3ca-c60cffc899a8 · outbound

This paper cites Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

Reference 10

Resolution
metadata mismatch
doi, observed 2026-06-26T14:49:32.413823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:ee67ee5eb486eaeb568918dc0998c00de31600c833a7e604cb0b7930ad0e15fc

Observation 51591905-effb-41a2-a969-7b89c7156504 · outbound

This paper cites Stable reinforcement learning with autoencoders for tactile and visual data.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Stable reinforcement learning with autoencoders for tactile and visual data

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-26T14:49:32.430508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:5ecf66a8d613b89dfbe6944c9dd70bcf35a14ceaafe6b549c24ea8c7cea93e25

Observation 60ca205f-8503-408c-8794-e2fdb5edb566 · outbound

This paper cites Decoupling dynamics and reward for transfer learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Decoupling dynamics and reward for transfer learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:04214b6a6505aed42b19fe98b62c1b5ae441cae87e95a02569bcfc92e3164da6

Observation f584e372-1751-4173-928d-e2c29e7e249c · outbound

This paper cites Bellemare.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Bellemare

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b45e01361682c8f4ab9548aa540e3da930f29cc7d198f32e45c593a4b326a39e

Observation 2a896039-a8cd-4dcc-b3de-8ff12ace07f2 · outbound

This paper cites Jha, Toshisada Mariyama, and Daniel Nikovski.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Jha, Toshisada Mariyama, and Daniel Nikovski

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:bb8435e31d46e646ce1fae37b736fb0f6e521fe1b0beea5bb431cc3762e99b2c

Observation 7099ecbf-4a99-40fd-bd70-e66041e3a368 · outbound

This paper cites an unresolved cited work.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:7f1c3f17fa88ee8894bfd3f9ed6a10571206e0e09a092fe794eb42c32aa31369

Observation 75dc07cf-ccef-40e9-8d4f-2f95001ce8d9 · outbound

This paper cites Bootstrap latent-predictive representations for multitask reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Bootstrap latent-predictive representations for multitask reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:2ef59243ca9cf8de32e1de4ebb297a0c8e81f89b6cab5270f8b408acd3249735

Observation 0810e75c-fb20-4721-9423-0c3e0bdedc8c · outbound

This paper cites Devon Hjelm, Aaron C.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Devon Hjelm, Aaron C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:122a30b72140746b9823cb642680756d61477487313c2b2791e3a2a0f56f2b5e

Observation a18538d6-d2fc-432c-a04b-030da0c899c7 · outbound

This paper cites URLhttps://openreview.net/forum?id=uCQfPZwRaUu.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning URLhttps://openreview.net/forum?id=uCQfPZwRaUu

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b7e35591b1f84c9b6816dc4a23e7bbbc95f03be957b628fe0fa45a709897478b

Observation 56bcb789-c91f-400d-aeec-0a16338bdfff · outbound

This paper cites Smith, Shixiang Gu, Doina Precup, and David Meger.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Smith, Shixiang Gu, Doina Precup, and David Meger

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:85034929f53ae9f5d3e2dadaeb7e39f297be12b3f17dcf9a872c3125d82d5d75

Observation 7bd3d6f4-c10a-41f3-9174-fb88b820ce9c · outbound

This paper cites Towards general-purpose model-free reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Towards general-purpose model-free reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:2b6af1d2836425ac8d33a9f261abc44ed5669f5e7b0231a5f8b60c3b024a6563

Observation 5e6401d3-d5f0-4d1a-81b7-dc2be5bb9b94 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:c4f183e12035e7fd0474fd375d65fb51cdd71ef232add1224591d6159984692b

Observation 2e20d636-adf6-4bd4-9e48-255d15caefe7 · outbound

This paper cites an unresolved cited work.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:e25a95dbc81b7b7fc7e8fadfe342e2304c474cf4ede6a103f86178b67e95dbaa

Observation f69f27f8-0109-4c9d-94be-6f7ed488a389 · outbound

This paper cites Layer Normalization.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Layer Normalization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.633607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:de12e3b2710239d24af06b9c540763e2ffe8b708e4be54237be2cdfa14bbed7a

Observation f4769078-8a5e-44c0-9e16-c06db084e6b9 · outbound

This paper cites Gomes, and Kilian Q.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Gomes, and Kilian Q

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:3d185ce739ea22661184ff5b4e1e80cf2359bac3ed1b59983f868211d0ee03d3

Observation 12312637-a44f-4785-b5ae-4be9ba5ee939 · outbound

This paper cites Spectral normalisation for deep reinforcement learning: An optimisation perspective.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Spectral normalisation for deep reinforcement learning: An optimisation perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:6036e6ae257502b53bfbdabb1c1b838a137aa28ebc1e948626782f73493618de

Observation e5ad6fa1-7c80-4d97-9115-433a0f76a642 · outbound

This paper cites Normaliza- tion enhances generalization in visual reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Normaliza- tion enhances generalization in visual reinforcement learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-26T14:49:32.422512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:186c8acd2b3296c17b4708e9c731b8ef864ba762ad2ce2bf176d7fb6a17ee7c5

Observation c1ee5e32-4f32-4d38-981c-94a0cb97a9a6 · outbound

This paper cites Image augmentation is all you need: Regular- izing deep reinforcement learning from pixels.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Image augmentation is all you need: Regular- izing deep reinforcement learning from pixels

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:c9f526bd7776953013dbea3f07de1010d59562822dd6a7cd3adae9c7ade3a002

Observation de6c538e-f278-41d9-bc83-a03dcefb9670 · outbound

This paper cites Mastering visual continu- ous control: Improved data-augmented reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Mastering visual continu- ous control: Improved data-augmented reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:563b03792cf49af0c29ab396965d7fc150ac95bec421201345a90602a8e313a7

Observation b6acbbef-2764-4ca7-887f-f1232dbdddbd · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.J.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Stable-baselines3: Reliable reinforcement learning implementations.J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:1706a5e939c2992dfeb92fdc9662e71c61bf15b575a296409e1d7abd22d25aef

Observation 622075b8-7b93-45c7-af33-d1770de7bdf2 · outbound

This paper cites Bridging state and history representations: Understanding self-predictive RL.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Bridging state and history representations: Understanding self-predictive RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:d53057d692cb8adceeb6196350a74a78c0fe424b7cd7b903ba902c2fba1f28f3

Observation a2d42450-a42c-4a3e-89ca-fa30ce881879 · outbound

This paper cites When does self-prediction help? understanding auxiliary tasks in reinforcement learning.RLJ, 4: 1567–1597, 2024.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning When does self-prediction help? understanding auxiliary tasks in reinforcement learning.RLJ, 4: 1567–1597, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:2b5cfe522410e5730add49f355a0da1015b74aa037ba23d070e0dedc67a767d9

Observation 95ec457b-7db3-4286-bdc8-45f8c7a71231 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:160a24d76cdf495755434b9cf9b69ef93258b7cc23ba2912452457d3ab126cdd

Observation 7adb7113-0e16-4b3f-bff8-09e97adae3db · outbound

This paper cites Balis, Gianluca De Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Markus Krimmel, Arjun KG, Rodrigo Perez-Vicente, J.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Balis, Gianluca De Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Markus Krimmel, Arjun KG, Rodrigo Perez-Vicente, J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:92ad4ae2d76e2f1c61befcd28b6c3610fc8be886cd79d4ef275ec48d86cc613d

Observation 44155838-b25e-4acb-a51d-10ede4b4108b · outbound

This paper cites DeepMind Control Suite.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning DeepMind Control Suite

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:19:37.628282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:948d31d0e5995aa07d6707c0b7ecad723d71b9e68b6551e3bf314d45402b9c66

Observation 84a96cdc-990e-4866-b683-c629c7c42091 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Adam: A Method for Stochastic Optimization

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.630932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:eecb8b3790ecdb8c862fc531475e19c73f4561571c7d1aee07b6788c6275c612

Observation 071d88f1-eded-41e6-82b6-328420f18bb9 · outbound

This paper cites Rainbow: Combining Improvements in Deep Reinforcement Learning , booktitle =.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Rainbow: Combining Improvements in Deep Reinforcement Learning , booktitle =

Reference 36

Resolution
verified exact
doi, observed 2026-06-26T14:49:32.411582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:d1ce2cff762374cdf499fac8166e5496e20f110d3ddc9cc5917fb02f32e2487a

Observation 3688046c-7490-455d-bc22-04cc85973582 · outbound

This paper cites Robust estimation of a location parameter.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Robust estimation of a location parameter

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:910182d2e335f41f8d0ea07fb679b12d70aced4f8475bc39d84a8996bf54971d

Observation c339a442-fc22-4ead-bd4c-eb7edef6a65a · outbound

This paper cites An equivalence between loss functions and non-uniform sampling in experience replay.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning An equivalence between loss functions and non-uniform sampling in experience replay

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:a8f44ec9cb3b0bfc7a9eae8194e97fc4da12f568282a5634f86e34eaf6e04348

Observation b21413f5-0cd0-4580-8395-8088a950a018 · outbound

This paper cites Ried- miller.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Ried- miller

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:fb33b7b11c46c73007bb3d0fe178eeba52f1d39f5c3a19f502083afd8221880b

Observation 28f94e41-b9b7-4174-8c84-f488652a871f · outbound

This paper cites Gomes, and Kilian Q.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Gomes, and Kilian Q

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:974977393035bea4b7ebfb368426600e9fe7658b2fc2e4b3e00f7fe53709b77e

Observation 0fea252d-e049-48a9-b141-5d9fec2ac46e · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:c497d0f41fa13f5ba6c10c2b68a978abeea1e1bbb9b1d57473cdaa187dfed311

Observation f2e322c0-12de-4360-babf-36c21c9be447 · outbound

This paper cites Discrete off-policy policy gradient using continuous relaxations.Unpublished.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Discrete off-policy policy gradient using continuous relaxations.Unpublished

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:d1bc06cde986f79a2fb848ba27dec10158063844977e4299d288f195eaa4c3a9

Observation ddd0e042-f6e3-47b4-b20e-e8f84cf56dd1 · outbound

This paper cites Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:01b0db671839705d63bda54125951447168776a961c7c081738a0906ae6b5a75

Observation acbd6fdf-f805-496b-a1a1-cb8dc53a7631 · outbound

This paper cites Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b52faff0f296e5dd4ca8a49f6aa0ccb79d14d076c39a41cfe4ac6761ad2efa10

Observation 5a1548d1-08b3-48c1-a79c-59d42c78956d · outbound

This paper cites Mujoco: A physics en- gine for model-based control, in: 2012 IEEE/RSJ International Con- ference on Intelligent Robots and Systems, IEEE.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Mujoco: A physics en- gine for model-based control, in: 2012 IEEE/RSJ International Con- ference on Intelligent Robots and Systems, IEEE

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-26T14:49:32.427750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:d6d274410a835e911dd3724bf58d500b41fe9f21bdda70210433a42442f50777

Observation 72437488-486d-46f5-8d05-29da19bc37b1 · outbound

This paper cites Spectral normalization for generative adversarial networks.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Spectral normalization for generative adversarial networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:04d3dad0089513c482f0f3abc4cbc47376d3fffd7058738b5dc6a549ec0b4f5b

Observation c1b9bb78-1254-4a8a-b3af-9744eea62d22 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Dueling network architectures for deep reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:1a0c1920397306d7a2b8785c0f82c777cd79f7b7060ac0c9738b2a63efda0528

Observation 1a174d8b-c689-41a3-800c-befdff8f9d78 · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning JAX: composable transformations of Python+NumPy programs, 2018

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:b23250d0b41408a90074a11abb1f858b756568bab94d2d0df9cf8fb1d79ed70d

Observation 310ac5da-b01e-46eb-ad1f-15f36e7a63e7 · outbound

This paper cites Fast and accurate deep network learning by exponential linear units (elus).

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Fast and accurate deep network learning by exponential linear units (elus)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:0ef44aca87e44d025c3e66ab5a5521bb2f23e70c1d1bc5b5c9a94ee2c753f90a

Observation 703738fe-1ce1-4ab0-b0e2-72bdfc00c7af · outbound

This paper cites an unresolved cited work.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:8b08c8fa95ab2193468d2d303da32a2331e07f599bb06eb1f52b6648625c460b

Observation b977c3a3-8490-4030-b852-1728b24c78b6 · outbound

This paper cites Deep sparse rectifier neural networks.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Deep sparse rectifier neural networks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-26T14:39:59.778081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:150cebe99cf3caa124d92fe2ad413def0fe6d1a66bc9ffe7232c26c0c3eb5ba3

Observation 33f75801-431f-47ae-a5e9-74ad184df649 · outbound

This paper cites Dying ReLU and Initialization: Theory and Numerical Examples.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Dying ReLU and Initialization: Theory and Numerical Examples

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.636482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:ad9bf2f3010871fa9a491a2e6db4597429a68362605162c9b57150b0d594f114

Observation 5224b758-6602-47de-9c78-c57758a2bede · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Backpropagation applied to handwritten zip code recognition

Reference 53

Resolution
verified exact
doi, observed 2026-06-26T14:49:32.425002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:d43be890a276488d9d21f37ada8630e6fa05c6a8a02a2affdb615c9fa4d15aaa

Observation 2f4399f8-fcfb-48f2-bceb-8ce79c2f1cae · outbound

This paper cites Gradient-based learning applied to document recognition.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Gradient-based learning applied to document recognition

Reference 54

Resolution
verified exact
doi, observed 2026-06-26T14:49:32.419739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:73168a259b032bac8201fb560c059f75bad170175129e25c2928daae00628b6d

Observation 2d3e7178-465e-409f-bce0-316b9ddc3f9d · outbound

This paper cites Zeiler, Dilip Krishnan, Graham W.

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning Zeiler, Dilip Krishnan, Graham W

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-26T14:49:32.417551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:39:59.778081Z digest=sha256:8fea91c5a873203659dd9101d5e4980021f7546918b19c698ca6ea142b19c7ac

Pith citing papers

Observation 0f5a31f3-017b-48bf-9a5a-4e7962ce4f1c · inbound

ProDVI: Programmatic Dynamics Priors for Value Network Initialization cites this paper.

ProDVI: Programmatic Dynamics Priors for Value Network Initialization NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:25:38.442865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T19:25:38.157587Z digest=sha256:5f8dfa6cfb3531535ece022c1c1fd3af1624c2f9d399b712fb3d2baedf7fb76b