Pith. sign in

Paper Citation Record · LEDGER

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.14019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14019 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:42:59.404858Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c84053-f80e-4b8f-84c1-453c19f3cc2e · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.123943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.123943Z digest=sha256:24c9101027787008039b953fb50ee462e1fbc884ded2be58ec9fb94204ae3126

Observation ecdca55a-7bb2-4742-aaef-1a34505bbcd0 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.744959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.128254Z digest=sha256:cee2d048b5fc11666986c59b35e15281f54ed20f1881efc72d75c3d0f15d6c27

Observation 381d1499-5fe6-4378-a685-f75d8aa4fea9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.132112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.132112Z digest=sha256:4e3ac6928461c489fb70cd85d4abd8bd3cebcc27db6004104f5ac01a88843fc4

Observation 24cd254b-6099-4857-9682-86463961ca40 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Dota 2 with Large Scale Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.136421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.136421Z digest=sha256:ada2ee3664ddb4b76e3a36f9858855796875c60e6766ed2356e526610313321e

Observation 24576388-8192-482d-9d48-c44095a59020 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.734943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.140358Z digest=sha256:79ad0890241408b087af09448aede0cdd15cd56ed9e9be64b95a0e9b6b593e27

Observation fb9eaa60-77f6-4d12-9569-17dc64c5a582 · outbound

This paper cites & V an Roy, B.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & V an Roy, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.724610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.144006Z digest=sha256:ff065d46786abc130f5d47e20136bb8fc3e97925e6835152373488255a16149b

Observation 99d7a4e1-9531-47c0-9a6c-36dc7f0e74f7 · outbound

This paper cites & Williams, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Williams, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.714635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.147694Z digest=sha256:385e58ab523167906e7e859fe86a6dc262d5e5ee0abaef4a4a348345b6967fec

Observation 83f9c814-d6f3-4b7a-a427-2955f06f1e62 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.151129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.151129Z digest=sha256:07cad4a71da220ab27316b565c93ade80ebc00cac52944901e6f79a24eb031d6

Observation 6aca79d4-1d02-4420-aef9-2238ca940b1d · outbound

This paper cites & Sutton, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Sutton, R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.697294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.154534Z digest=sha256:bdf960d15e89f177752eb49a9ab5c74319e480302764531253c031ea5620dca5

Observation f87e1e2e-1a36-4a68-acc0-825f438294a9 · outbound

This paper cites Algorithms for reinforcement learning (Springer nature, 2022).

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Algorithms for reinforcement learning (Springer nature, 2022)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.687679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.157947Z digest=sha256:9b2992d194ee3f2b7683b817d2253ef6e47b6d42d3f4812451d1a7cc9b45928b

Observation 7a9d7340-920e-4246-9be3-e5a2dc69fb58 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.676781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.161517Z digest=sha256:90639664466d7db7f270fc1ac76bdd94076b1d0e21bbd1926add6e26965da847

Observation 32a26af1-7437-4fd2-8094-5ba61c633612 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.666649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.164937Z digest=sha256:07321ff72adc12d5cede5081ab47e0a814915ec50aef36d06afb797a16414a09

Observation 9e936308-dedb-45db-bfe0-127856906239 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.657054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.168279Z digest=sha256:01fb34c9a0c28d5160ebe8da5724cf8326f6cf60c882c80621de14a1d4a76379

Observation cec608e5-bf70-4f60-aa52-e78108fc2ecd · outbound

This paper cites & Silver, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Silver, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.646626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.175545Z digest=sha256:7fac6092e7fcb1cbb3d7eda917b91cc5e520c74c19ce8a1208dc5a81361dd8cd

Observation b2894724-a0e9-4ba2-9528-9381dd00129d · outbound

This paper cites Error bounds for approximate policy iteration.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Error bounds for approximate policy iteration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.636542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.178859Z digest=sha256:2a9de45983239c3382514ff98535290f4c698a85853231202b606b82cc5f4eb9

Observation 5391d55f-7134-4852-851d-8a29e11141e1 · outbound

This paper cites Prioritized Experience Replay.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Prioritized Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.182279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.182279Z digest=sha256:9fb903fc5d34411380845c9edbd430fc9aaf40af94e5d8199037664b52dc2575

Observation 76a94055-7c46-4536-8752-8daad7f98cb2 · outbound

This paper cites & Pilarski, P.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Pilarski, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.625890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.185862Z digest=sha256:f0a865b75c37e67ec22c1f5217780c38f81845d7aface96b7f854916939eb9d1

Observation 9b2e9293-a66c-4e0a-a9e8-bf86c7793da1 · outbound

This paper cites & Wen, Z.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Wen, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.189256Z digest=sha256:bc2cd9e1ff488d64b92e18d292415da0cbbec4d9e48c32f8bc86acf071380cd9

Observation 4e491de7-3065-4406-9cef-a33d81355653 · outbound

This paper cites How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.192879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.192879Z digest=sha256:553aeadadc22796ba889b59751f365e10a15d081e73d6ab189a6ad5cd62f46f5

Observation 3e042fc0-aab3-41b1-8c54-9faf9ea65bcd · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.604663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.196488Z digest=sha256:502a6eeed61f3a09f981c5adb943e03ad8029b13e4f32478df5e70f9a81d1425

Observation 15c7bbf3-4bd7-4f00-991f-56d3cfba2e47 · outbound

This paper cites Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.592919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.199768Z digest=sha256:d801c3e4807a8e37bcab0721151ed2402ff1b9c9a3c982828fdc72a3dd132ecf

Observation 25bf3420-43d4-42aa-a0ea-de2cbe694ad5 · outbound

This paper cites Discovering hierarchy in reinforcement learnin g with hexq.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Discovering hierarchy in reinforcement learnin g with hexq

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.580832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.203160Z digest=sha256:70f587eb87818fdefea53960998d663405204c93fd051c58a1292b1b67da719d

Observation 0de67471-4fdf-4303-af8b-bdaa098d74ad · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.570128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.206660Z digest=sha256:2da31c9798828280dff241d2f81bc1d89dc5ba0011475b0abd11657a0d860e1a

Observation 2ee35c95-28e3-4b5c-9f77-d23cdeac66fd · outbound

This paper cites & Shimkin, N.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Shimkin, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.558865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.209835Z digest=sha256:4f0e7457dad4b40ee251a86c2accb6aa3179d4fec1b1fd59a5ea9851ccb10600

Observation f9f7ba4d-e162-4c3c-94cd-6b9179d2d483 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.547755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.213035Z digest=sha256:c5b0d2acea914ae7fa0a54c81fb84c2af23c7ea702653f6750c87e1afdb83a27

Observation 08db37ee-5cc7-4356-ac02-a81ee85b08a1 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.216225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.216225Z digest=sha256:16efc8ec39f9249751702afd955f86e8d38a86c8ccafdc82f12d48a43ce1d2f0

Observation 107ff731-acc9-4c2d-a8ee-4c29de087b79 · outbound

This paper cites Human-level control through deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Human-level control through deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.529899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.219702Z digest=sha256:15cecd5e4daef1e530b7139663a2120575eeaa32d6d2b8ccb467389397db0b56

Observation 4e8acf22-2e3c-486b-a08d-25bec08e83d0 · outbound

This paper cites A markovian decision process.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition A markovian decision process

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.223137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.223137Z digest=sha256:2c836d75463043693839721c6b8f72b010a774cb956aaf89418216a2bcaf38fb

Observation a7040e14-bba6-4992-a7ad-7b1cee181c07 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.226498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.226498Z digest=sha256:ad230243f821546fb67a8c7d89d090b057457203a68cce6cb8cbf9f5df5e98b8

Observation e20b45e4-f732-4a9c-b065-32c7bdef6a0b · outbound

This paper cites S., McAllester, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., McAllester, D

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.513143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.229784Z digest=sha256:021f62348e4235aaa0c75338c7ae48cdd2df094f00197416c9348de828a5f7d1

Observation 5ebf1381-c522-491a-b65c-3f76c5b8b3e9 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.503178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.232992Z digest=sha256:734cec79f77bbf5af8e6626380436ec193311c038edfa4c44257d7325e7cc803

Observation ff5edea6-8a09-4782-aeef-b1888b8d85ab · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Asynchronous methods for deep reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.236178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.236178Z digest=sha256:9e79f643da8229a31bb7ea2c2bb7563442b0f3982fae3d7d544da7b02285c5b0

Observation f2de305b-222f-454b-8927-97ea947fc5d7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.239471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.239471Z digest=sha256:b3aab6cffefb9a3a9ff9a4f8d5500d314bac5764a0f4e618cbba072cb9c4dfbf

Observation 870b958d-5b33-4066-97d1-fdef6dbd1879 · outbound

This paper cites S., Barto, A.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., Barto, A

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.243122Z digest=sha256:1c6a4acde59a49724c3a347c14f4afb1fbb16063363e51fc9bf06dced053ff33

Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · outbound

This paper cites Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.246831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.246831Z digest=sha256:dbae755b2e99b314cccc69dfca044b7e6f23ff96d4ab5560957da4f3b8d106e3

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · outbound

This paper cites Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:9202dff5c5da1fc61a9cf89b4a422ffbb8841af5a2ab82299e0c4c10714cf98d

Pith citing papers

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:9202dff5c5da1fc61a9cf89b4a422ffbb8841af5a2ab82299e0c4c10714cf98d