Pith. sign in

Paper Citation Record · LEDGER

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.14019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14019 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:42:59.404858Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c84053-f80e-4b8f-84c1-453c19f3cc2e · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.123943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.123943Z digest=sha256:24c9101027787008039b953fb50ee462e1fbc884ded2be58ec9fb94204ae3126

Observation ecdca55a-7bb2-4742-aaef-1a34505bbcd0 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.744959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.128254Z digest=sha256:c23497f00d397baeb6e1efdbe95f1356cf612a54ab8e2447a801f30b1da2a725

Observation 381d1499-5fe6-4378-a685-f75d8aa4fea9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.132112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.132112Z digest=sha256:4e3ac6928461c489fb70cd85d4abd8bd3cebcc27db6004104f5ac01a88843fc4

Observation 24cd254b-6099-4857-9682-86463961ca40 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Dota 2 with Large Scale Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.136421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.136421Z digest=sha256:ada2ee3664ddb4b76e3a36f9858855796875c60e6766ed2356e526610313321e

Observation 24576388-8192-482d-9d48-c44095a59020 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.734943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.140358Z digest=sha256:a8134445877fc452738f866e853d062bb0e243a3df83f678a521b997e1892ac0

Observation fb9eaa60-77f6-4d12-9569-17dc64c5a582 · outbound

This paper cites & V an Roy, B.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & V an Roy, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.724610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.144006Z digest=sha256:547ce895a1642770760d4219a9ac919306acb38dc6e2642f638a91ea23344e50

Observation 99d7a4e1-9531-47c0-9a6c-36dc7f0e74f7 · outbound

This paper cites & Williams, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Williams, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.714635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.147694Z digest=sha256:57140de85dda973aebd1c6831de0e52f5fae7547caa602a507346d8b4f25dac3

Observation 83f9c814-d6f3-4b7a-a427-2955f06f1e62 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.151129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.151129Z digest=sha256:07cad4a71da220ab27316b565c93ade80ebc00cac52944901e6f79a24eb031d6

Observation 6aca79d4-1d02-4420-aef9-2238ca940b1d · outbound

This paper cites & Sutton, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Sutton, R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.697294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.154534Z digest=sha256:78833a3af4c0446acee82d37e832b148586c7c8956526ae4f10b5b0fe589ab34

Observation f87e1e2e-1a36-4a68-acc0-825f438294a9 · outbound

This paper cites Algorithms for reinforcement learning (Springer nature, 2022).

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Algorithms for reinforcement learning (Springer nature, 2022)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.687679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.157947Z digest=sha256:b76a702c54ba11ec5b62132da66db8938d1837d0a13c334ba6ca8803f2e9cdc2

Observation 7a9d7340-920e-4246-9be3-e5a2dc69fb58 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.676781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.161517Z digest=sha256:c5a138f3ecd3713550e621af17be4e0b6bac8f41ff3676b4b57ee214ff1f5bc9

Observation 32a26af1-7437-4fd2-8094-5ba61c633612 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.666649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.164937Z digest=sha256:b7c89b258001d1bf6204a9a52ef3eca964523a3e3db0cbff2fe0dcc24922d4a6

Observation 9e936308-dedb-45db-bfe0-127856906239 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.657054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.168279Z digest=sha256:79d116b3258bd226ee115faf09716b6c7be393d6eacc7c1bfa932714ed739801

Observation cec608e5-bf70-4f60-aa52-e78108fc2ecd · outbound

This paper cites & Silver, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Silver, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.646626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.175545Z digest=sha256:6e3aa0458c76bec4b8303f17e177efa23116282ec055dd673d2a2e71d734c4e0

Observation b2894724-a0e9-4ba2-9528-9381dd00129d · outbound

This paper cites Error bounds for approximate policy iteration.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Error bounds for approximate policy iteration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.636542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.178859Z digest=sha256:a729b51246ddc29c54cf2ade64938b1506c507c97939922407216560c869e5b6

Observation 5391d55f-7134-4852-851d-8a29e11141e1 · outbound

This paper cites Prioritized Experience Replay.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Prioritized Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.182279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.182279Z digest=sha256:9fb903fc5d34411380845c9edbd430fc9aaf40af94e5d8199037664b52dc2575

Observation 76a94055-7c46-4536-8752-8daad7f98cb2 · outbound

This paper cites & Pilarski, P.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Pilarski, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.625890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.185862Z digest=sha256:df40a743bf03e90d7c95b755f994d6697f7a10cad5414fc4a7ea025ba78edc2f

Observation 9b2e9293-a66c-4e0a-a9e8-bf86c7793da1 · outbound

This paper cites & Wen, Z.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Wen, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.189256Z digest=sha256:9bfaeded2eee9d15c59aa82e4d499b2e801feb71fc7c7936e02249e2ff2ee9be

Observation 4e491de7-3065-4406-9cef-a33d81355653 · outbound

This paper cites How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.192879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.192879Z digest=sha256:553aeadadc22796ba889b59751f365e10a15d081e73d6ab189a6ad5cd62f46f5

Observation 3e042fc0-aab3-41b1-8c54-9faf9ea65bcd · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.604663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.196488Z digest=sha256:d44000a8fb5b8347e45c859c095f50425083c6dcc2289017ab5002968f79f7e1

Observation 15c7bbf3-4bd7-4f00-991f-56d3cfba2e47 · outbound

This paper cites Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.592919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.199768Z digest=sha256:aee78ab651d994aedd6f7a6e6429e56c47f7b1e85cebd1eed2c05d5493799dda

Observation 25bf3420-43d4-42aa-a0ea-de2cbe694ad5 · outbound

This paper cites Discovering hierarchy in reinforcement learnin g with hexq.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Discovering hierarchy in reinforcement learnin g with hexq

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.580832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.203160Z digest=sha256:df91f44c9cad762c310dfda7454df42a254e822bc38867b68469e3442416e3a1

Observation 0de67471-4fdf-4303-af8b-bdaa098d74ad · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.570128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.206660Z digest=sha256:1e92eae8c504a19b6df67777abee149a8234cdf31d64fd6e10db308b028989cb

Observation 2ee35c95-28e3-4b5c-9f77-d23cdeac66fd · outbound

This paper cites & Shimkin, N.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Shimkin, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.558865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.209835Z digest=sha256:c983ae0108cdcf8b60929e945cfabb6b966ee96e7b470318bf83dc91a18526b0

Observation f9f7ba4d-e162-4c3c-94cd-6b9179d2d483 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.547755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.213035Z digest=sha256:44fcfc3b4dfce66f14e2822ae0bb3ebda6868629f94c56fd76c45361f1ae181a

Observation 08db37ee-5cc7-4356-ac02-a81ee85b08a1 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.216225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.216225Z digest=sha256:16efc8ec39f9249751702afd955f86e8d38a86c8ccafdc82f12d48a43ce1d2f0

Observation 107ff731-acc9-4c2d-a8ee-4c29de087b79 · outbound

This paper cites Human-level control through deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Human-level control through deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.529899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.219702Z digest=sha256:08d409b0e9ebc02460cd329a3c895a354cfc774acbe42bf00c39911aedb902fe

Observation 4e8acf22-2e3c-486b-a08d-25bec08e83d0 · outbound

This paper cites A markovian decision process.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition A markovian decision process

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.223137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.223137Z digest=sha256:2c836d75463043693839721c6b8f72b010a774cb956aaf89418216a2bcaf38fb

Observation a7040e14-bba6-4992-a7ad-7b1cee181c07 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.226498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.226498Z digest=sha256:ad230243f821546fb67a8c7d89d090b057457203a68cce6cb8cbf9f5df5e98b8

Observation e20b45e4-f732-4a9c-b065-32c7bdef6a0b · outbound

This paper cites S., McAllester, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., McAllester, D

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.513143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.229784Z digest=sha256:aa22fd777ac9039145947da2bcbb6e5de9286a81a7e53ebf6238a1557448ca43

Observation 5ebf1381-c522-491a-b65c-3f76c5b8b3e9 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.503178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.232992Z digest=sha256:0c8440707adace1ff915b064b6ecb2c214332d483887bbc5d15e803aa6538fd8

Observation ff5edea6-8a09-4782-aeef-b1888b8d85ab · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Asynchronous methods for deep reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.236178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.236178Z digest=sha256:9e79f643da8229a31bb7ea2c2bb7563442b0f3982fae3d7d544da7b02285c5b0

Observation f2de305b-222f-454b-8927-97ea947fc5d7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.239471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.239471Z digest=sha256:b3aab6cffefb9a3a9ff9a4f8d5500d314bac5764a0f4e618cbba072cb9c4dfbf

Observation 870b958d-5b33-4066-97d1-fdef6dbd1879 · outbound

This paper cites S., Barto, A.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., Barto, A

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.243122Z digest=sha256:1c6a4acde59a49724c3a347c14f4afb1fbb16063363e51fc9bf06dced053ff33

Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · outbound

This paper cites Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.246831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.246831Z digest=sha256:dbae755b2e99b314cccc69dfca044b7e6f23ff96d4ab5560957da4f3b8d106e3

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · outbound

This paper cites Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:23cbe05412b8215b4b26c37e5f3f3975f109e34e23aeee7f63d0ab280e0c6d0a

Pith citing papers

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:23cbe05412b8215b4b26c37e5f3f3975f109e34e23aeee7f63d0ab280e0c6d0a