Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.06937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06937 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:56:51.152652Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be150886-de80-4762-a6a8-42644b6bad6c · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.596693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.042565Z digest=sha256:fd86c4b8397572b8230feffa00a308f3e83ceb7744d4912566f421e8a02b32f9

Observation 93c12656-9030-4f74-8e1f-8cba2be7ec58 · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks G., Naddaf, Y., Veness, J., and Bowling, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.587835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.046887Z digest=sha256:8b712d66a1b1ee577af63c5ad6c977d8fda6d825261af8ae29788e21a489c569

Observation d14898ed-0568-4795-a409-d931006dbf8c · outbound

This paper cites OpenAI Gym.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.050601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.050601Z digest=sha256:5c6c0b7675f0e096031563c54534b3d1e85aff6763d58e3c18b2ee59314276a8

Observation cbe57015-1b5f-430c-b418-b5b2bd07c762 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.579136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.055134Z digest=sha256:d23a65df99ad7d0fa7a1b5c51fd62d7aaadbe954be1ac1289bcdc28b9badba44

Observation d2fb5b3c-811c-4c25-bd61-bc90e63e0b63 · outbound

This paper cites Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.059025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.059025Z digest=sha256:16fb249a830c6bfe908fc1055ca0251c606bfffa75bd07ca0ca085cc22b5ca37

Observation 4da91d0c-c520-49ec-9ea2-c098833d9a8e · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.062911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.062911Z digest=sha256:8da781d96800a75a13ff3e54bd1bcedaabf006664cb8a3b0723e1f7414e8c64b

Observation e803cdbb-79e9-457f-8208-ec2546712ee9 · outbound

This paper cites and Petrik, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Petrik, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.564149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.067089Z digest=sha256:e6045443f4dda149f5b95a1326624e9cd7b7f2a1e48341cb14a1d9b1f44d76a2

Observation cbebf486-942b-40e4-9eb7-2f4cadf725a9 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.070550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.070550Z digest=sha256:37a8410f3ab5406b3de809983ed4e1c54332bbc7075d849f6438cc4b72d15262

Observation 3577ed12-96a2-4ede-afe8-fbe2a3f4b295 · outbound

This paper cites RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.074552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.074552Z digest=sha256:11a23a54a16be3e48f82ca4c292a9485e81b6525d1cecf24102edab443f0b5b3

Observation 5c5ea7ba-f326-4ded-8cca-75798e2af797 · outbound

This paper cites Continuous control with deep reinforcement learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Continuous control with deep reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.078679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.078679Z digest=sha256:9eb80605cc1d63d829e5c08a27cc2e2c4a61b4d6aeaa234fa614232a0504237b

Observation 0bc6b9d6-5972-481e-acdd-4019b2ad8d9a · outbound

This paper cites Average-Reward Reinforcement Learning with Trust Region Methods.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Average-Reward Reinforcement Learning with Trust Region Methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.082564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.082564Z digest=sha256:392ad67e2a47a7d9dcd6f4c529ed01eeb7ac1e3191b66801ce8d87e3f47364f6

Observation 835823a5-df91-487b-a6f5-ab970128cb0b · outbound

This paper cites A., Veness, J., Bellemare, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks A., Veness, J., Bellemare, M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.087603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.087603Z digest=sha256:dff7d4bc5ebb72b280b61a2be4a44cc68b5cc3145624c50526d156a2edf9c551

Observation 9f9ef73d-43e4-433f-be92-dde05275b5d1 · outbound

This paper cites Reward Centering.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Reward Centering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.090784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.090784Z digest=sha256:f085cc183331d368e27b4bc6fc4695f1ea8fee29949962204637dbec1467751b

Observation 9822b3bd-4d1c-4ac9-84f5-6723d589abaa · outbound

This paper cites Jelly Bean World: A Testbed for Never-Ending Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Jelly Bean World: A Testbed for Never-Ending Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.094911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.094911Z digest=sha256:098bd320963297628f52c56888d52fcde42c1fe4e6997d7a8ceb8c988d9faf9d

Observation b3a91d80-5572-4fc1-a42c-e842cb5e85ba · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.549244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.098506Z digest=sha256:177f41f644a4145860b7021b8a25f5492bb8729713b03014d394f24cc93419db

Observation b351cda6-3d10-4448-9656-c9523f29e2ad · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.538881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.101851Z digest=sha256:9c74aba7971e0b8d09b57048a9912c4bf0a5c160ada2fe6d3e64f8b954e7eb1c

Observation 519c400b-9380-42fe-89b6-375b81f9ba7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.104986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.104986Z digest=sha256:35efaa274114aab9075d3c2a73cef36bcd182947aa505c326f4aeadfdd278267

Observation f765d528-5c76-4c55-892f-ea08522312a1 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.528853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.108332Z digest=sha256:22c16eabf1eb36fa2057deda5fa46ccd0fc4fb86ce4bc6fe90cc6fa1fa869798

Observation 7d38b37d-f827-444a-9ddb-b1a33c365145 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.518642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.111420Z digest=sha256:25d43f8832844ae2a62a420f906e57a84c89ba17ec4ffa22438730696656f9b9

Observation fb175955-ecca-48b7-8b67-7ef70bbd0fc0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.114743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.114743Z digest=sha256:69e5cbd5045ae581020049e832e2d47fd013b25648969bd55f0fcced087efc36

Observation a58ef605-33d2-43b5-97bb-51781700d1d0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.117972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.117972Z digest=sha256:3f3a7d0485c268249af3465cc1116282bd5666c3f3aa6cd2da6e8e4b6813b3ff

Observation 3fcd1410-718a-4941-88da-70582473c74f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.121085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.121085Z digest=sha256:61c2ec8ba987b6fed26751ca49bf028d200c4b961732a2d99f434a1b04d30da2

Observation 32099d3d-34fd-4afb-9ad3-424c62f33400 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.498595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.124304Z digest=sha256:50103ef658ba9293c9c35470cb1592ce6de1712e401b31bbcc7c3f2112a1b0dc

Observation c768aa35-d9c8-4e8c-b19c-23112d861c4f · outbound

This paper cites and Ross, K.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Ross, K

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.488242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.127557Z digest=sha256:895bbdd897ff0c322b674aeced0bc1577ceb01504994abff4f972049e2442c68

Observation 05517a7e-0d4f-4860-88f3-f56349d0ec11 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.130679Z digest=sha256:7e1d7541c0ed418c7f1aef183b35d781d0588e7b51815acc4899ce02442e9dd3

Observation 3c0fb13f-2a7f-4cd3-ae6c-34f461c50e48 · outbound

This paper cites The Ingredients of Real-World Robotic Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks The Ingredients of Real-World Robotic Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.133757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.133757Z digest=sha256:b5c8fb1b811bb75b4cb83187c64ed2a17afc3456fa76489674da978f4c4666a8

Observation 47f1ddd2-023b-4c39-8c16-9f6ae5e4777a · outbound

This paper cites Pearl: A Production-ready Reinforcement Learning Agent.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Pearl: A Production-ready Reinforcement Learning Agent

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:56:51.352949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.137298Z digest=sha256:422c0f7d5434bec92697aa4561de7455a2bb4e184567ff4cb2d41a28d3829d39

Observation e88cee22-f087-413d-a8ff-26293f0d462b · outbound

This paper cites write newline.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.141171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.141171Z digest=sha256:ff936d1541e8f5519dd7afbad521eaafc60d0de313354cdcb93e15c47106207e

Observation 432d02ac-23d8-4d74-b0e4-b4f2566d4be1 · outbound

This paper cites @esa (Ref.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.145296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.145296Z digest=sha256:56dea87ad2c246265db4703fec6e0617aeb69549a20e5c211f34eae49c78b9c1

Observation 0c86f820-0284-474a-ba05-541e0913f5ac · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.149172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.149172Z digest=sha256:95c3042f6d7fde8b8a1e9614ec5ff5edfc7592e92923567dcfbe974a0d191a68

Observation 152e618d-2879-4b74-bba0-39abe64e30a6 · outbound

This paper cites yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:56:51.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.152652Z digest=sha256:5eace5a200e228a672097f6d825a2337fc70e957b447ccd145dd29ceaf72f685

Pith citing papers

No inbound Pith citation observations are available.