Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.06937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06937 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:56:51.152652Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be150886-de80-4762-a6a8-42644b6bad6c · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.596693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.042565Z digest=sha256:e96acd183af0acbf62933f55e7356e52f07ef8df156f8700381cd47c3847bbc3

Observation 93c12656-9030-4f74-8e1f-8cba2be7ec58 · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks G., Naddaf, Y., Veness, J., and Bowling, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.587835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.046887Z digest=sha256:4ee9a56ff5e51328de708fca13f2efebd547ec32e3afd30e825505dd955ce28e

Observation d14898ed-0568-4795-a409-d931006dbf8c · outbound

This paper cites OpenAI Gym.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.050601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.050601Z digest=sha256:fb16e27ef54bc6acc9d3f572777605cccd88ff27d3877ad3d7d923562eae17cb

Observation cbe57015-1b5f-430c-b418-b5b2bd07c762 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.579136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.055134Z digest=sha256:47b0ab06fb192a2d18a8167fe45223d8542f67b8219fb437c1095cab7d42dfb9

Observation d2fb5b3c-811c-4c25-bd61-bc90e63e0b63 · outbound

This paper cites Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.059025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.059025Z digest=sha256:b32575a887c991e671975a38042c2a4a5cb295f1d425ac13e8f7c622c41f21ab

Observation 4da91d0c-c520-49ec-9ea2-c098833d9a8e · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.062911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.062911Z digest=sha256:e61143fcb069cd35e89018191699b742178de927fd5e525e103882ede467b89f

Observation e803cdbb-79e9-457f-8208-ec2546712ee9 · outbound

This paper cites and Petrik, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Petrik, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.564149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.067089Z digest=sha256:86590077c385ca8b9709e9258558d3791b283ef863ebe1f5526337ada6dbaebd

Observation cbebf486-942b-40e4-9eb7-2f4cadf725a9 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.070550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.070550Z digest=sha256:9969864bff2afd610dffea4e843e319841212a4413ad090d834a94a8bf8328e3

Observation 3577ed12-96a2-4ede-afe8-fbe2a3f4b295 · outbound

This paper cites RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.074552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.074552Z digest=sha256:b680f09bf2f8121e96993869792f9379fe43fe2aa1687bde6556003cb0e67d8f

Observation 5c5ea7ba-f326-4ded-8cca-75798e2af797 · outbound

This paper cites Continuous control with deep reinforcement learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Continuous control with deep reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.078679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.078679Z digest=sha256:c4c535be613dccc4a4a95990cc5f91c9e056677bda41f63d8ee510a3a7c65727

Observation 0bc6b9d6-5972-481e-acdd-4019b2ad8d9a · outbound

This paper cites Average-Reward Reinforcement Learning with Trust Region Methods.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Average-Reward Reinforcement Learning with Trust Region Methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.082564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.082564Z digest=sha256:961281efe8e6fd3802f76864476729b01875f430f82e488ab62a31ff788a4d07

Observation 835823a5-df91-487b-a6f5-ab970128cb0b · outbound

This paper cites A., Veness, J., Bellemare, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks A., Veness, J., Bellemare, M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.087603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.087603Z digest=sha256:4dd66c39d3ab81f8f27d0c099623a79f0e161e01250b899fd4125e212d9bcde7

Observation 9f9ef73d-43e4-433f-be92-dde05275b5d1 · outbound

This paper cites Reward Centering.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Reward Centering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.090784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.090784Z digest=sha256:13e841cea0f767de925ab7f1c89913776cadd726b43feb1b4712866a88857edf

Observation 9822b3bd-4d1c-4ac9-84f5-6723d589abaa · outbound

This paper cites Jelly Bean World: A Testbed for Never-Ending Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Jelly Bean World: A Testbed for Never-Ending Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.094911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.094911Z digest=sha256:1006a4fcc7d1c803fa65ccf3c51d302c17795dca848c250e63713048bf6454e9

Observation b3a91d80-5572-4fc1-a42c-e842cb5e85ba · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.549244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.098506Z digest=sha256:1b67f0c376e58f97b03fafad9a88ed21d5e373e242ab76b9420443e87c095e37

Observation b351cda6-3d10-4448-9656-c9523f29e2ad · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.538881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.101851Z digest=sha256:1a7cbf95447de637867567fdbbf0e0cb00ba71371fdede2cabcd662262894830

Observation 519c400b-9380-42fe-89b6-375b81f9ba7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.104986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.104986Z digest=sha256:f58098d9c4c4caa486997336b82675761cebefb8488bb4c8379bd82b37f686bf

Observation f765d528-5c76-4c55-892f-ea08522312a1 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.528853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.108332Z digest=sha256:9677daa1f96f917b8731700c25a723e2d5d1118e5bf4e538a02ac3e2fe04f9f1

Observation 7d38b37d-f827-444a-9ddb-b1a33c365145 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.518642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.111420Z digest=sha256:ba0a0294e34eae9db427f60a3a6779559146ad48255f883701fa5c9c646e9a92

Observation fb175955-ecca-48b7-8b67-7ef70bbd0fc0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.114743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.114743Z digest=sha256:2907d7f99eaeabef7c41cfb6e43729193dab2b420d7f9654e1fb1d8f6ba83d2b

Observation a58ef605-33d2-43b5-97bb-51781700d1d0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.117972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.117972Z digest=sha256:6b55c7fadb099e238131e848e9acb0c253ae78f38cbddf73968b82cb05bc3515

Observation 3fcd1410-718a-4941-88da-70582473c74f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.121085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.121085Z digest=sha256:55240aa8c94b05e8ce4fd38d7192cd4fd968a77599ac0fb54756efaec9563f00

Observation 32099d3d-34fd-4afb-9ad3-424c62f33400 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.498595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.124304Z digest=sha256:b2037fc5c5591e5113f86aabe60d649f40ec0a316d736c9c94496cfc2ebe9ed5

Observation c768aa35-d9c8-4e8c-b19c-23112d861c4f · outbound

This paper cites and Ross, K.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Ross, K

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.488242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.127557Z digest=sha256:7a6a6c7aa49be74b8d6ba68e10d1e82eff467761bd0538344e01500dd1a1d9db

Observation 05517a7e-0d4f-4860-88f3-f56349d0ec11 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.130679Z digest=sha256:77f01d0ffff1bc285169723d4e019bc98939819cf71aa8265362c8c9f72f7914

Observation 3c0fb13f-2a7f-4cd3-ae6c-34f461c50e48 · outbound

This paper cites The Ingredients of Real-World Robotic Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks The Ingredients of Real-World Robotic Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.133757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.133757Z digest=sha256:fb933fea38723009671871f2ab0db47ae6cf728b7805ab075d1abfcdffc41cd7

Observation 47f1ddd2-023b-4c39-8c16-9f6ae5e4777a · outbound

This paper cites Pearl: A Production-ready Reinforcement Learning Agent.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Pearl: A Production-ready Reinforcement Learning Agent

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:56:51.352949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.137298Z digest=sha256:29a74ea2f3e08a6093ce4a046a788ac568e6ccb9c3911249b0825f049c43753e

Observation e88cee22-f087-413d-a8ff-26293f0d462b · outbound

This paper cites write newline.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.141171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.141171Z digest=sha256:2ebd552a42df05b34e906a6f43b5c31bd80e43af8d86be3a8a6da18fad0015ae

Observation 432d02ac-23d8-4d74-b0e4-b4f2566d4be1 · outbound

This paper cites @esa (Ref.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.145296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.145296Z digest=sha256:7b8ac32af396c990a8f10c59a8979f0259f405866a450ab72f70636fa11ad0f9

Observation 0c86f820-0284-474a-ba05-541e0913f5ac · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.149172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.149172Z digest=sha256:2665d6fd54b0064765403d108abf520bf361238b66de0e8be5659a3070f70205

Observation 152e618d-2879-4b74-bba0-39abe64e30a6 · outbound

This paper cites yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:56:51.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.152652Z digest=sha256:cf03d38853018a0bd8948bf7084ed27932b9f541121adde238ff59515a3da72f

Pith citing papers

No inbound Pith citation observations are available.