Pith. sign in

Paper Citation Record · LEDGER

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.09523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09523 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82aaab77-0338-46e0-bb06-92e787f30d06 · outbound

This paper cites write newline.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:44.735602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:44.735602Z digest=sha256:209f807ec0d3ca0c09bb8e7f649aa6d006ed80671525da17f76b87ed7048e4d8

Observation a2eb5b45-b5dc-4f01-b32a-fa21b259404f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:51.426569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.855222Z digest=sha256:5e9d440610a4323257413798f338c337242b24bb251c5389ba854acc8a64159f

Observation 49aac7d9-f0e4-4fd7-be52-3ff2dcc6abd5 · outbound

This paper cites Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:51.113175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.976043Z digest=sha256:70528772b8e11ac075d4cb8c79090f9f87c8ea95863885dac8f552e80385ceaa

Observation 88c31345-0586-4a52-b1f5-a7038a19a76d · outbound

This paper cites Dynamic Programming.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dynamic Programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:45.055174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:45.055174Z digest=sha256:05d752270f912dc4cae9e590b61c18e60b0b0a30f2c7252e9ff26cdb87b9a64c

Observation 7450fe15-d8bf-428c-9bda-665639b9468a · outbound

This paper cites Bertsekas and John N.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Bertsekas and John N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.904931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.192835Z digest=sha256:0d37ce5dbfa576dbe908546cc7b69e5cd06ad843b416f987504a488fffb21270

Observation fbe8b06b-4a0d-4413-a3e5-73037c103b4f · outbound

This paper cites Multistep Credit Assignment in Deep Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Multistep Credit Assignment in Deep Reinforcement Learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.775974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.304714Z digest=sha256:4821717e90ab2ff39cfbe85edb97b47c9cf185709983baedd884c131705a8dab

Observation 49119f6f-aff3-4fbb-8560-2f69162cbcff · outbound

This paper cites ChainerRL : A deep reinforcement learning library.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values ChainerRL : A deep reinforcement learning library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.632827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.368932Z digest=sha256:17444b98dc5d2018602bfa12fcc72a5288b40865fb6a7a79a0c1b3467aa969f7

Observation 7dfa3418-b4fe-4d31-83b9-87ac0e446569 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:50.491388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.486551Z digest=sha256:47d2d83f2763f3b935d16a4022a50e8040702c7beddbc8630f345a8108d4feea

Observation d742ebfd-07fc-4337-90c9-25fd85e9f059 · outbound

This paper cites Kingma and Jimmy Ba.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Kingma and Jimmy Ba

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.360877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.604642Z digest=sha256:fedef8c19d8f69df5f1cff13e38ede97ea04b98bc03f8b7d1e3204f50bd398c5

Observation 58a7a71b-afa8-4ac7-82b8-9c0fbb5dfb7a · outbound

This paper cites Orr, and Klaus-Robert M \"u ller.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Orr, and Klaus-Robert M \"u ller

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.213833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.697423Z digest=sha256:0bec2bfd1c12fd4a5116f245306bf5a1e6a5e80c2dd3aca1061dac72b739058c

Observation 9a3e49df-698b-4a37-8b30-74688ca8b8b5 · outbound

This paper cites Rusu, Joel Veness, Marc G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rusu, Joel Veness, Marc G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.064504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.818776Z digest=sha256:be6543856d9b18d70a66242d1b864736a714c0479628c332c6740be2039c4b85

Observation c9259be3-76b0-4b10-a6a0-c393a0c51e4d · outbound

This paper cites Towards model-free RL algorithms that scale well with unstructured data.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Towards model-free RL algorithms that scale well with unstructured data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:01:47.519541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.917099Z digest=sha256:580cdcbc87e70946c5bfb46e5a9e7ff44a6974a54a49e0f0cba6d5b7407ecbdf

Observation 03ae75ac-1be6-466a-bae3-d767da002530 · outbound

This paper cites Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.937782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.020479Z digest=sha256:24b1e627cd25bce01d0ab08ad559d9ddd955aa434b0305edc18bf2c268797630

Observation b6d1f872-1467-4ef7-83ea-07d58b4a3290 · outbound

This paper cites Rummery and Mahesan Niranjan.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rummery and Mahesan Niranjan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.834689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.115279Z digest=sha256:df3c6d0973400f6b3e3602007a08fc7cb9709ebd10f3b3f31d0c4923a27f96ca

Observation 49ad9446-5ee9-4e3d-b989-db0a0616c173 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.703269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.207382Z digest=sha256:634c1dffadca3d82198c5ade57dd22d5629b24ba859c23d16c6459139acd6e4b

Observation 8cff4409-b4e7-4c88-bd2e-fb44c043a8b3 · outbound

This paper cites Littman, and Csaba Szepesv \'a ri.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Littman, and Csaba Szepesv \'a ri

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.575152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.290297Z digest=sha256:f5f2cd94983fbe1cefe3ce887b6ac1c0458adae7a55b54d4f8afe52a94f8b77f

Observation 58931b2f-9137-4d8e-a2c7-a1c07d8fb3ce · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.377144Z digest=sha256:abc68b50f4ae23abb9f5aacafa29bcb0298baa0927947577767fc637d8d9695f

Observation 90ab8c3d-19a9-4e2f-8f9d-8f470a925902 · outbound

This paper cites Sutton and Andrew G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sutton and Andrew G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.289687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.468196Z digest=sha256:8f60cef722c751bed733879b92d530e49fa96a92d6b853bfd8a9ecb046828eb3

Observation 380e6ba9-f74a-4281-accb-a82f3045e7c2 · outbound

This paper cites VA-learning as a more efficient alternative to Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values VA-learning as a more efficient alternative to Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.135279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.563594Z digest=sha256:e16d528a4caee36c5c2bc5f0fc6a99bc8c4026836e8536cae538eafea6f75af4

Observation 41bad30e-a29f-44fd-993e-b5942ea11e2f · outbound

This paper cites Double Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Double Q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.638334Z digest=sha256:374cedf782ef2ac2ad0505b171827542a57048127b88da6d3620deac68b3e082

Observation 2413f22d-5b75-47ec-b6f8-1940c20fd601 · outbound

This paper cites Insights in Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Insights in Reinforcement Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.840394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.723290Z digest=sha256:7ca615fa2888a424f0b42baa1723a89c09975e966bd6d298ba0d681a2255fb43

Observation 344721d1-d006-4cd5-be36-51e75d1f6521 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dueling network architectures for deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.710073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.799484Z digest=sha256:afbb59e7b4d4e4b9fce659e9bbf3291006f8ad5850a69a2b5e88f188d82a24c9

Observation 3bc160fe-f2bf-477d-8e61-9713e3fe5b9d · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.572660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.815493Z digest=sha256:6333e9119a4d03a7b66d2c55ee80ee5545d29e7631e19636fcce9f19690733aa

Observation 15cdb5f2-f4e9-43b5-91fa-fc55735cdd2f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.406295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.827140Z digest=sha256:51e44275d70a0fd5186624a05f86fa74e1c22079d9968ea91f4c58baac20b2f5

Observation a390f672-f32c-42d0-aa0d-86a298d88ebc · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.101275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.894642Z digest=sha256:f5f716f7d3373190b0193fa505adc3c22c2acbd56137b96c295c8e947fed99d6

Observation d490c7c4-2992-4a36-8d87-4706443e9f22 · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:47.826751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:01:47.067674Z digest=sha256:f6d971f5c664fb7bc597116ec88b171b2c1e2a2e0c3cd1c616c8513b12804db3

Observation 85b62dba-ea0f-4746-a7cd-3f2f496889b6 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:47.191696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:47.191696Z digest=sha256:9423f96c8e800084bfebe0c270524ee15d7c6aa64ba9e14df884e5bbb134c161

Pith citing papers

No inbound Pith citation observations are available.