Pith. sign in

Paper Citation Record · LEDGER

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2501.19254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19254 v4

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:58:30.177720Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:34:05.838427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T22:34:09.275036Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36993655-962a-4d50-9df2-2f42138b16d9 · outbound

This paper cites write newline.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:58:29.390261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:58:29.390261Z digest=sha256:449166230c96106a22ea99bd13c908c1b62d412ad437e4bedff2721707b44eac

Observation ca0f78ec-4fd1-4ba3-829f-25503527166e · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:32.332945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.496381Z digest=sha256:8fde494a0ec77bb112d7c9fc9e35277521ddfaaa1a136ed23858686b5658be8d

Observation 0cf194cd-20f1-4293-8500-e43a9462d597 · outbound

This paper cites First-order methods in optimization.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set First-order methods in optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T20:58:29.530301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:58:29.530301Z digest=sha256:8971e7b4394172db81f0c0e204c3f77a6995864ddd4fa26cd2ce7a91cd4850c8

Observation 21f23488-367d-429e-9ee2-49433e08f2d6 · outbound

This paper cites A Markovian decision process.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set A Markovian decision process

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.302788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.535413Z digest=sha256:bbe9e59b245f2ecde840fd938807bbba2929319c8a34bdd46ebd6c6e769be51a

Observation 74714a0b-bcff-446d-b52c-f2664339746b · outbound

This paper cites Adaptive Algorithms and Stochastic Approximations.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Adaptive Algorithms and Stochastic Approximations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.282169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.543242Z digest=sha256:fdd875df473ad871532366d29a8fa680720e5a95a2343e165aacf49f71ef87b8

Observation 22e9fad4-e53d-4e1d-b15b-d1738b0bf16d · outbound

This paper cites Stochastic approximation: a dynamical systems viewpoint.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Stochastic approximation: a dynamical systems viewpoint

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.261902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.549783Z digest=sha256:6669473c9663e9753363bc0c8c17d8a92b1aec52ec5372bb2a145d3f68df847e

Observation be507291-899f-4fe1-b39d-81fa98b0a64e · outbound

This paper cites The ODE method for asymptotic statistics in stochastic approximation and reinforcement learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set The ODE method for asymptotic statistics in stochastic approximation and reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.243181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.556360Z digest=sha256:66e6e3be7f64528de3e6add3cade8569218409fbd986a01437f0ece900b9c0b3

Observation f523b574-2638-46fe-823f-fb200063030b · outbound

This paper cites D., and Wang, Z.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set D., and Wang, Z

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.225964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.563163Z digest=sha256:224f5f7b2438cec6c9ac4a0885b3128684a717501f8556266393e14ee66afb31

Observation 8e655a9c-0dd4-4419-8e84-2830d3ec41d4 · outbound

This paper cites S., and Santos, P.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set S., and Santos, P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.191334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.570563Z digest=sha256:9a0309ee3bdddf024838ed0b70c5494283d94ee703ca322d9438bda7c39e8664

Observation e4df219c-019c-4c79-a391-498b63dda8e8 · outbound

This paper cites and Zhao, L.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Zhao, L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:32.099214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.576330Z digest=sha256:75c031e4b3f5d970d319864541fc708caba28724a803ee1ee8c5a06956bc4020

Observation b3f740b2-af26-4a5d-95e9-fae02cfed326 · outbound

This paper cites T., Shakkottai, S., and Shanmugam, K.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set T., Shakkottai, S., and Shanmugam, K

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.964475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.581772Z digest=sha256:7a776f9b00859af64911e34da4d6d1971da817c51ad6c5dad2a8badf5e94aff2

Observation 94dade05-3c28-47fe-912b-ebcba9b1d1dd · outbound

This paper cites T., Maguluri, S.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set T., Maguluri, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.828721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.586922Z digest=sha256:fa93865dd7343593b929c55f73ce57d4ea6bab0458318310c5fbac2e6dea4363

Observation 147eeda4-a667-4d21-b5d8-2e685798b24b · outbound

This paper cites P., and Maguluri, S.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set P., and Maguluri, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.806526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.592659Z digest=sha256:3767f8376ab3aec85dc6c5de8f4b57b2b432a8d19719f2f50e51afef4c35586e

Observation e3597ec1-2087-432a-ba1d-6caf3d63e160 · outbound

This paper cites T., and Zubeldia, M.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set T., and Zubeldia, M

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.783135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.597792Z digest=sha256:1f45fbd16e007be5bce588953254b50efcda2d087779c620784fd4bc250414c4

Observation 097a0c5c-d6c2-496a-b271-0a673294059e · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:31.757980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.603381Z digest=sha256:02ec04bd6fdb795932ea95d982f3eebfd0f3f13a7fdf434c96e27d9fc4004222

Observation e49b25aa-5768-4420-8d29-fe28c3fa3505 · outbound

This paper cites Learning rates for Q-learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Learning rates for Q-learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.737817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.608375Z digest=sha256:041d8279442bb75ac00df2c19d7329a052ca3679832f04f204dcb2a3ea90c104

Observation 50e030d0-6e9a-437f-999a-63e28936ddd0 · outbound

This paper cites A theoretical analysis of deep Q-Learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set A theoretical analysis of deep Q-Learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.717377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.613242Z digest=sha256:68bc5a9f17aac7a429928869636c92e3c4ccb417722e4a238d1a2de1f9142195

Observation 30024e98-be1f-4070-b1a6-65b4403d43f6 · outbound

This paper cites and Thoppe, G.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Thoppe, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.699415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.618420Z digest=sha256:93728c71789c50f24fed6a05b727f741ef2d9e3d598fbe9c917acd313fc8c9da

Observation cf449a89-630a-4e74-89f9-e759d6ed5071 · outbound

This paper cites Boundedness of iterates in Q-learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Boundedness of iterates in Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.679866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.623231Z digest=sha256:165e41ee48b839bf96bb7556a52a3c86a023fdd15c30d877aff39ab44ffc4f5a

Observation ab2de571-7c9a-41f0-8742-7e7a1f2c8191 · outbound

This paper cites and Donghwan, L.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Donghwan, L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.659209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.629140Z digest=sha256:ff20d126c9cc8b823c64bcfcc81f0d618b874cc5066a29cdf65f3ee4009f741c

Observation 2abc488c-6815-4fd9-acd7-2fc770dcfc56 · outbound

This paper cites Convergence of stochastic iterative dynamic programming algorithms.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Convergence of stochastic iterative dynamic programming algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.641631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.635962Z digest=sha256:46571539527d31b1f8bcdc905c26cd3170e828cfe4751dcabdf58debed5830fd

Observation af605be1-789c-4f4c-bcb7-40ab2e86f5a4 · outbound

This paper cites Exponential hardness of reinforcement learning with linear function approximation.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Exponential hardness of reinforcement learning with linear function approximation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.622732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.641535Z digest=sha256:69cb31c79ffe3ef6a5e0f9cf08ced39f8365af5ce97d89aaf4bbb44b54edf0e0

Observation ac00abf5-f772-4bcd-bbd4-8dacaeef6af5 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:31.542458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.646863Z digest=sha256:d60938ecd91539f38548ea815c5139e91910125008f64d1dac233cab8403deea

Observation 6af5b5c5-1147-434d-96e9-48c17f3b1a0d · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:31.436348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.652892Z digest=sha256:a52e0b9fb2133c3d70d79883aac7784c2fe432bcfea9c55ff3f7b54f16ff4f35

Observation 61401d30-56de-455b-8acb-013d1d23378c · outbound

This paper cites and Yin, G.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Yin, G

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.275608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.658470Z digest=sha256:0f981b699b8bf57525ba1dc06626ec49d19ad9145b35ddcbb0272bb0e52b5401

Observation 86a32984-15b8-4566-b662-518aa055615a · outbound

This paper cites and He, N.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and He, N

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.239096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.666504Z digest=sha256:1f1dfd895516e9b39ab9ecfede829ce4ddf47d1f6908ca05ffbb6c52d2dc3dd6

Observation 735f8e06-d9d0-4a82-b6ca-5b5f14708115 · outbound

This paper cites Is Q-learning minimax optimal? a tight sample complexity analysis.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Is Q-learning minimax optimal? a tight sample complexity analysis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.217145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.672969Z digest=sha256:8e344e9fdde18d07df1d98d9cb6b136db7db3806128b7dc064bffa56b3d0d75a

Observation 586ba490-444e-4b60-9e33-2ac17e9460ac · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:31.197787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.678067Z digest=sha256:1d0211c8774d4f749ea83075e1204a75b21968eca9c3ed88ba2dc926a2305a60

Observation 3d0236e0-685c-43c6-80c3-a0d3a2f45425 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:31.170326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.684277Z digest=sha256:2755eed2fd083a486db4ee8b2611ca545ccb62b2e865ea1043bd4dfb96dbcacb

Observation 91d28fff-0ec5-4729-87ef-b5883f8c0ce9 · outbound

This paper cites The ODE method for stochastic approximation and reinforcement learning with markovian noise.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set The ODE method for stochastic approximation and reinforcement learning with markovian noise

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.145455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.689700Z digest=sha256:7c47e49d86bdd5b2a521bd0690b7908f80722f5b3794f6afc1ffa89ccbbae364

Observation d353bbdf-7db6-42d9-8054-da0396cfdaec · outbound

This paper cites and Tsitsiklis, J.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Tsitsiklis, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.124280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.694822Z digest=sha256:fac306f8b4c8dcfe971be0d6949b1bd2b56341a422e138211cd741b99dfccfe6

Observation 25d3d80d-462a-4c8b-bf90-2d5b3f0ff733 · outbound

This paper cites S., Meyn, S.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set S., Meyn, S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.098987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.699755Z digest=sha256:e4045660812cee792ce604f7915ac0c57e7fa4cebcad5e192b1fc85f8eb86b3b

Observation 2abb70ce-03c6-4999-ae83-c49b3790b1c1 · outbound

This paper cites The projected bellman equation in reinforcement learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set The projected bellman equation in reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:31.076321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.730600Z digest=sha256:197dfcf8a603cd5a4b0cfe19390391871666a1f6e470444f90fc5302cd865333

Observation 93ecb8cc-cd3f-466f-ba2b-79f345a76951 · outbound

This paper cites A., Veness, J., Bellemare, M.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set A., Veness, J., Bellemare, M

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T20:58:29.797237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:58:29.797237Z digest=sha256:f55aee2363802443728c09f7a3745ca5c23dd83380d07a1e2fd98aa61e93a732

Observation 3135a3b8-9ff9-43c0-881c-a461fe61bd84 · outbound

This paper cites and Gharesifard, B.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Gharesifard, B

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.914278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.829769Z digest=sha256:c70f7bedbf1496f10dd4cba1130f1464da05faa63f18475d1d1f6f9da6751998

Observation 10df2676-522a-4a07-a883-b352416d928a · outbound

This paper cites Almost sure convergence rates and concentration of stochastic approximation and reinforcement learning with markovian noise.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Almost sure convergence rates and concentration of stochastic approximation and reinforcement learning with markovian noise

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.797136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.888380Z digest=sha256:e33988b104b3a9aaf991107958b3fbec5569183e840ad81c09c1281487fd7a04

Observation 91a537ef-454f-41f3-957b-7f6f17892b11 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:30.694645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.922180Z digest=sha256:aa90ec020b4d7d6e3662c9a8d5ed0fa9347c72f7d8afdfc47fc6c2f549ad5ffd

Observation 48b790ea-a766-46f5-a630-5c0803fe35b9 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:30.659539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.971995Z digest=sha256:65aa0880acf520f27de1597ae38f2313dade5c1e722ea3c25d4d71b560425a49

Observation b7f7c519-4037-4712-91a1-5ea044dc45e4 · outbound

This paper cites The asymptotic convergence-rate of Q-learning.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set The asymptotic convergence-rate of Q-learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.642271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:29.997454Z digest=sha256:32fa6eee0ddf279a76c1318efcb80e5a4e0c072db5a7238c0afe1462240547a5

Observation a44c37db-d70e-4e71-8cc1-574dd8c7b253 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:30.624256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.078504Z digest=sha256:93a3d315b337b791cf9c21afed19a326aa3be2b28db3b39a279e5927875733d5

Observation db4ebd1a-0cb3-4be6-9bd7-ae655731f9a9 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:30.606665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.087584Z digest=sha256:e65f83ce6f188ad516f7cf13895e48279cc135ef90b5090b0db46571bd1de996

Observation f2c0bf0b-ba49-463e-b799-fb40a67a6088 · outbound

This paper cites an unresolved cited work.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:58:30.588206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.100942Z digest=sha256:bbe59615cbcf545f85aef94a0d08ab845eb0596e3eed4b250d57d68694d94b74

Observation 25726b4e-0b1b-4379-8dab-857216be19cd · outbound

This paper cites A finite-time analysis of two time-scale actor-critic methods.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set A finite-time analysis of two time-scale actor-critic methods

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.568048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.109221Z digest=sha256:1ed751d67db34b1fa3341ee32763603a11b40860c5af2c7dc542fd02ea55f045

Observation 97c06f47-f27e-4604-af1e-454b8882815f · outbound

This paper cites and Gu, Q.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Gu, Q

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.549325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.115028Z digest=sha256:482f7b1d5d64a0ef95ebb1c2c239ecbd13c00f186aad727ed6574ac74e34893c

Observation a6697c8a-c688-44d6-86b2-06108e35c5ec · outbound

This paper cites Provably convergent two-timescale off-policy actor-critic with function approximation.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Provably convergent two-timescale off-policy actor-critic with function approximation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.532477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.122062Z digest=sha256:bfdd27a0487b90e686bcc5b20066bfbd5879cb617f944c66af7bce2fddd68804

Observation a3138edf-0121-4d6d-b436-10662e8720bf · outbound

This paper cites Breaking the deadly triad with a target network.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Breaking the deadly triad with a target network

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.476178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.131992Z digest=sha256:6b33e95179d2b38d8bf7c61f698d65b932b77d3b4cf49e94337687d45cb2912a

Observation c163ce65-0e5d-47f4-8744-bba20c400e72 · outbound

This paper cites Global optimality and finite sample analysis of softmax off-policy actor critic under state distribution mismatch.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Global optimality and finite sample analysis of softmax off-policy actor critic under state distribution mismatch

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.343119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.142792Z digest=sha256:9a1ba01e503288c80482ab7ef04cccb70a9d1dac3f80f0aba419fb51105b0d2e

Observation ebf40e84-ab56-4b01-9aad-90382fe31612 · outbound

This paper cites T., and Laroche, R.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set T., and Laroche, R

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.315884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.152536Z digest=sha256:0443ae161558bfe280ac0ecb823a3f39e89a27fe38d4fde7511b53eced0c13f3

Observation 843744d8-6dd9-4332-8b3d-d730a85110ab · outbound

This paper cites On the convergence and sample complexity analysis of deep q-networks with -greedy exploration.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set On the convergence and sample complexity analysis of deep q-networks with -greedy exploration

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.292104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.159097Z digest=sha256:6138f30d66e1ba7cdcd02d679be38ef5a10caf712711dcf2da0def58921251ca

Observation 5cd36632-0109-4d69-a2ed-d0fda9ef87fe · outbound

This paper cites and Xie, Q.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set and Xie, Q

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.269383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.171211Z digest=sha256:a8055c4bf1a492cd5fa7ba996e8267510afe3f69e5d14c48d48d1d2506ea8875

Observation 44efca40-5a04-470f-8d0e-0b08cb185131 · outbound

This paper cites Finite-sample analysis for SARSA with linear function approximation.

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set Finite-sample analysis for SARSA with linear function approximation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:58:30.241760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-09T20:58:30.177720Z digest=sha256:36ab83391e7edb636bcdfd46f230b0c59226a083a1fc90353178164898367757

Pith citing papers

Observation b42121ae-070b-4859-85e6-37be898e1cdc · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.163269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:dcdc707ec3ea399d7cf1efdc613e39c271ca654cfa6ea3e3a0d8ddf04af8996b

Observation 94bc028b-1789-43f7-b855-2690b0f69a4c · inbound

A Switching System Theory of Q-Learning with Linear Function Approximation cites this paper.

A Switching System Theory of Q-Learning with Linear Function Approximation Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:17:22.850655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T06:16:23.809510Z digest=sha256:054df245cd2753f11bbd5cc9d93f181d8cfc1a4dec9fffc916a5c1023c372aa9

Observation 827a4c9f-7183-4d24-a336-a71cb2bad00d · inbound

A Switching System Theory of Q-Learning with Linear Function Approximation cites this paper.

A Switching System Theory of Q-Learning with Linear Function Approximation Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:34:09.281243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T22:34:05.838427Z digest=sha256:f6cf756581467f755a9e5e97d5c99baec91276b4177a68031c690809ed65f726