Pith. sign in

Paper Citation Record · LEDGER

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.09523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09523 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82aaab77-0338-46e0-bb06-92e787f30d06 · outbound

This paper cites write newline.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:44.735602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:44.735602Z digest=sha256:209f807ec0d3ca0c09bb8e7f649aa6d006ed80671525da17f76b87ed7048e4d8

Observation a2eb5b45-b5dc-4f01-b32a-fa21b259404f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:51.426569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.855222Z digest=sha256:5f704c4bacf51406e82bc42386e84a96c8f4a78eb8e8e2213606975cf20984a1

Observation 49aac7d9-f0e4-4fd7-be52-3ff2dcc6abd5 · outbound

This paper cites Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:51.113175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.976043Z digest=sha256:9af6669529033a4d5b5c3d2fdedb27d2fc888eda7ce82e960140996a3923c1b0

Observation 88c31345-0586-4a52-b1f5-a7038a19a76d · outbound

This paper cites Dynamic Programming.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dynamic Programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:45.055174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:45.055174Z digest=sha256:05d752270f912dc4cae9e590b61c18e60b0b0a30f2c7252e9ff26cdb87b9a64c

Observation 7450fe15-d8bf-428c-9bda-665639b9468a · outbound

This paper cites Bertsekas and John N.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Bertsekas and John N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.904931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.192835Z digest=sha256:2c627bbe544624c735024cfaaa8d0a5f340d1de035f5a528dc82a0cbb1a501b1

Observation fbe8b06b-4a0d-4413-a3e5-73037c103b4f · outbound

This paper cites Multistep Credit Assignment in Deep Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Multistep Credit Assignment in Deep Reinforcement Learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.775974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.304714Z digest=sha256:bf23d1db8bddfdab666acc8b3bcba14b776b7bbdced3e68ad9d8425b03c985e1

Observation 49119f6f-aff3-4fbb-8560-2f69162cbcff · outbound

This paper cites ChainerRL : A deep reinforcement learning library.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values ChainerRL : A deep reinforcement learning library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.632827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.368932Z digest=sha256:4377bb1c196426a3e98c77cbe201a716b7faf2aeab39ae2b8f3e0172e23f6e2e

Observation 7dfa3418-b4fe-4d31-83b9-87ac0e446569 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:50.491388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.486551Z digest=sha256:f8696771ecd6089e279b253429f1e340fdbfe39b002e39294f0d934fd2d3cebd

Observation d742ebfd-07fc-4337-90c9-25fd85e9f059 · outbound

This paper cites Kingma and Jimmy Ba.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Kingma and Jimmy Ba

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.360877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.604642Z digest=sha256:85afeee82cf6f35e67384f04a68d33bf456103b49f40d0712ea678ad45b42cfc

Observation 58a7a71b-afa8-4ac7-82b8-9c0fbb5dfb7a · outbound

This paper cites Orr, and Klaus-Robert M \"u ller.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Orr, and Klaus-Robert M \"u ller

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.213833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.697423Z digest=sha256:4c43cf505fbdaf014bedb8f2c3097c068b80c6383f33954ad6e1b62312e54b05

Observation 9a3e49df-698b-4a37-8b30-74688ca8b8b5 · outbound

This paper cites Rusu, Joel Veness, Marc G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rusu, Joel Veness, Marc G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.064504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.818776Z digest=sha256:ec83bbc7d125350a809ebcb36cad232a1594dad18c32be21abef2b78e0b167ea

Observation c9259be3-76b0-4b10-a6a0-c393a0c51e4d · outbound

This paper cites Towards model-free RL algorithms that scale well with unstructured data.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Towards model-free RL algorithms that scale well with unstructured data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:01:47.519541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.917099Z digest=sha256:83c274273ab9da155f158591b3d7099ab395a221f2a9d13a3d0477fbd552d3d0

Observation 03ae75ac-1be6-466a-bae3-d767da002530 · outbound

This paper cites Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.937782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.020479Z digest=sha256:7260a7f98712058f6991afeeb121a40e0f7f1a1a4d9a2197fcf8bbe2c54defc6

Observation b6d1f872-1467-4ef7-83ea-07d58b4a3290 · outbound

This paper cites Rummery and Mahesan Niranjan.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rummery and Mahesan Niranjan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.834689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.115279Z digest=sha256:8bf600a14d53c4e8421f18a7290d0af6056242a17dd94a5f6e81fc4c9f137191

Observation 49ad9446-5ee9-4e3d-b989-db0a0616c173 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.703269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.207382Z digest=sha256:6e0df4ab46b618e9dcc954af85cdbcd23afc101ef49d96a56b33cd76cc3cd21d

Observation 8cff4409-b4e7-4c88-bd2e-fb44c043a8b3 · outbound

This paper cites Littman, and Csaba Szepesv \'a ri.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Littman, and Csaba Szepesv \'a ri

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.575152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.290297Z digest=sha256:200d59fefc2fd9f7c01a9d92bc0b01ed34005214e03a036e2f6915465eda5232

Observation 58931b2f-9137-4d8e-a2c7-a1c07d8fb3ce · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.377144Z digest=sha256:0d4719ddfe3b79cf51ce8f4e88848bc199aa4e77f6b2f56b0aa820fb102f5efb

Observation 90ab8c3d-19a9-4e2f-8f9d-8f470a925902 · outbound

This paper cites Sutton and Andrew G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sutton and Andrew G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.289687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.468196Z digest=sha256:83cd6dc963445fb05548593ca9ea2b5fe775d5e232ce15fa2e48f53c71e61cf2

Observation 380e6ba9-f74a-4281-accb-a82f3045e7c2 · outbound

This paper cites VA-learning as a more efficient alternative to Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values VA-learning as a more efficient alternative to Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.135279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.563594Z digest=sha256:5c3470aa27ed2495b0a0bc0e808e5234e5ebd13ab20167ce69aeeadc1500f168

Observation 41bad30e-a29f-44fd-993e-b5942ea11e2f · outbound

This paper cites Double Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Double Q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.638334Z digest=sha256:a2253aacafb1cd3683f2b653a526383ba1a16adbb03fe6c30d11f436627b0f66

Observation 2413f22d-5b75-47ec-b6f8-1940c20fd601 · outbound

This paper cites Insights in Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Insights in Reinforcement Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.840394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.723290Z digest=sha256:1399ca7ee119c357f37e13c44aa76050b8cd39c7cab35ae0b812b6c525b7119c

Observation 344721d1-d006-4cd5-be36-51e75d1f6521 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dueling network architectures for deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.710073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.799484Z digest=sha256:b539d6caab0d28836a5cf3c25bb80668885ad66e6c933ff00179f10032b5f2f4

Observation 3bc160fe-f2bf-477d-8e61-9713e3fe5b9d · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.572660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.815493Z digest=sha256:af0a28352fe166a735065d03c3f9321ecbde63896bc45ab5ce67e0681e7815d7

Observation 15cdb5f2-f4e9-43b5-91fa-fc55735cdd2f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.406295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.827140Z digest=sha256:ab8d52276e9268b9521449280825badb7d82365759344b957d3eb62bd5ca9b8c

Observation a390f672-f32c-42d0-aa0d-86a298d88ebc · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.101275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.894642Z digest=sha256:36cc69511a7e7a1b86f8ad29aac38e082df9fd38f58c3169d7a41dd6a485139c

Observation d490c7c4-2992-4a36-8d87-4706443e9f22 · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:47.826751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:01:47.067674Z digest=sha256:68e6e9b9f13e527d576f8ed0f3084f97fd6286f1989792f28bf9c45cd3a15bdd

Observation 85b62dba-ea0f-4746-a7cd-3f2f496889b6 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:47.191696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:47.191696Z digest=sha256:9423f96c8e800084bfebe0c270524ee15d7c6aa64ba9e14df884e5bbb134c161

Pith citing papers

No inbound Pith citation observations are available.