Pith. sign in

Paper Citation Record · LEDGER

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2501.04870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04870 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:32:16.181011Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:10.654724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:49:14.996703Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0ec2c96-cbe6-47d6-965c-8c40046abc08 · outbound

This paper cites Offline Multi-task Transfer RL with Representational Penalization.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Offline Multi-task Transfer RL with Representational Penalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.109676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.109676Z digest=sha256:dbda4a9105e0413ec1069050244ea7be6adc4558d7a6bc273b480a52194838ab

Observation 2caaed7d-f951-4f6e-a72a-c1d286d793d3 · outbound

This paper cites After some algebra we get that ∥ bQp t − Q∗ agg t ∥2 nM,bPagg t ≤ ∥gp t − Q∗ agg t ∥2 nM,bPagg t + 2 nM nMX i=1 (byrwt−ki t,i − Q∗ agg t,i ) · ( bQp t,i − gp t,i).

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning After some algebra we get that ∥ bQp t − Q∗ agg t ∥2 nM,bPagg t ≤ ∥gp t − Q∗ agg t ∥2 nM,bPagg t + 2 nM nMX i=1 (byrwt−ki t,i − Q∗ agg t,i ) · ( bQp t,i − gp t,i)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.401105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.161325Z digest=sha256:aadd4c34b402ea5ff271ff654ae34ea3185a27c72bb1a742a4dcff246590c30d

Observation fcc1872c-9d52-44da-8afe-7255ce3a95a5 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.351905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.174866Z digest=sha256:0059eba771045636133b3f24e6ff27e1a43bae7c2697b5e5957bd82029e33e55

Observation 7575b9fe-bbf0-4a72-bd2f-5b7d84f48565 · outbound

This paper cites sub-Gaussian random variables with variance parameter σ.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning sub-Gaussian random variables with variance parameter σ

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.413812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.157923Z digest=sha256:44ec6327aba31628b4e106e86c40ff5d499423ac557c668166e607354833ba73

Observation 14068bac-a517-45fd-885d-c40053615633 · outbound

This paper cites copies of z, G be a b-uniformly-bounded function class satisfying log(N∞(ϵ, G, zn 1 )) ≤ v log ebn ϵ for some quantity v.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning copies of z, G be a b-uniformly-bounded function class satisfying log(N∞(ϵ, G, zn 1 )) ≤ v log ebn ϵ for some quantity v

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.562749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.154387Z digest=sha256:939e62d66729e3468cf0aa0e158bd8c48d3950c38fe28e6ba05ead7463f0c088

Observation 4c3e234b-a9cc-4f49-8492-0e8e88f5e2a4 · outbound

This paper cites Robust angle-based transfer learning in high dimensions.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Robust angle-based transfer learning in high dimensions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.122035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.122035Z digest=sha256:79c15179a9ea6110840ac5efb937a04cfaeddb563eb5a319602cea5288a82169

Observation 322b8543-e945-4058-bdad-7302491fa76e · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.388060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.164919Z digest=sha256:5481c0c6d4c6736841808af0ce7fde61ae180a77f24ae4e002377dd316068a37

Observation 222c81a9-84bc-42e8-a0cf-3c4dc036e6f7 · outbound

This paper cites We now use the peeling argument to extend to uniform r.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning We now use the peeling argument to extend to uniform r

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.376291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.168599Z digest=sha256:4b496da1f6b5f124fa4b269bafaed101499b3862715bc9c74d7e0fe79aa1fe01

Observation 13b356e3-f83d-492e-8913-7039e61e011e · outbound

This paper cites For g1, · · ·, gN being an ϵ-covering set of G, we claim that g2 1 − eg2, · · ·, g2 N − eg2 is an 2bϵ-covering set of ¯G.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning For g1, · · ·, gN being an ϵ-covering set of G, we claim that g2 1 − eg2, · · ·, g2 N − eg2 is an 2bϵ-covering set of ¯G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.364756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.171929Z digest=sha256:ad744265e39ac4c7c277ad054333f49fc462a8b23e5b3a661e5521d203eed82a

Observation 393c7740-b859-4d7e-a972-8aecdb99fde4 · outbound

This paper cites In our dataset, the mortality rate is 24.21% for female and 22.71% for male.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning In our dataset, the mortality rate is 24.21% for female and 22.71% for male

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.341161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.177942Z digest=sha256:46a606645943264153ce8fc8dc8871762c39a947e4330a456b3156a0d212fa94

Observation 7e83532c-c609-4c66-9c6d-a8b9e2d5d5f1 · outbound

This paper cites Figure 5 in Chen, Li & Jordan (2022) presents mortality rates of different lengths.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Figure 5 in Chen, Li & Jordan (2022) presents mortality rates of different lengths

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.329149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.181011Z digest=sha256:5144021cb28c66dd9b410b67ca31cef95889535694b419debab55ada16e1182a

Observation 26bbd280-67f7-439d-b3c7-ea930d053232 · outbound

This paper cites B., Davidian, M.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning B., Davidian, M

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.585575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.147497Z digest=sha256:310670ecf89cb33c984078e734b87db477ee410f31afd770bfcd1feffbd6a92e

Observation bba30a7c-7f58-49a8-8c84-50e975fb6d42 · outbound

This paper cites & Remlinger, C.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning & Remlinger, C

Reference 343

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.118274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.118274Z digest=sha256:6fd81c1c481c947e77afaac55c3464bc986ab47a8e54c5c8a23ca89fce183181

Observation a019a3d5-0491-4d7b-becc-8ee33a89f447 · outbound

This paper cites & Song, R.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning & Song, R

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.596606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.138500Z digest=sha256:7f469d68790e9a0a2f5222a3cc96b77cb4120e9bec318328a09dd76d5e10a8cd

Observation 523052cf-28b5-4f18-9179-f93475dfe160 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 651

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.630880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.114359Z digest=sha256:10ba9439b985b1c9dfb27cd9f3878f6f8e28f74b9dffe84c8f7f9dc7836847b9

Observation 65c1489b-4bfd-445a-b7c5-d722c61da4b5 · outbound

This paper cites Pseudo-Labeling for Kernel Ridge Regression under Covariate Shift.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Pseudo-Labeling for Kernel Ridge Regression under Covariate Shift

Reference 901

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.142623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.142623Z digest=sha256:3ee83b05452d7cf87a7e21af3e18ed42e63984689655b152bea1524f8acfd0bb

Observation 70a0fd0f-6126-4ca7-8351-939c8091aa40 · outbound

This paper cites (2012), Transfer in reinforcement learning: A framework and a survey, in ‘Reinforcement Learning’, Springer, pp.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning (2012), Transfer in reinforcement learning: A framework and a survey, in ‘Reinforcement Learning’, Springer, pp

Reference 1225

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.620368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.126456Z digest=sha256:34c8d621d4e53c933cf86893b05dc69932ee1bb6a39282eaf01e0c0f9d294202

Observation 6f06d249-26eb-4209-8c26-c81945a89489 · outbound

This paper cites Deep Transfer Q-Learning for Offline Non-Stationary Reinforcement Learning.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Deep Transfer Q-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1549

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.574542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.150696Z digest=sha256:676cf84798bb954701d9dd876e9d1e4eed1a4ffd88bbd0a6ec2a2968f8cd8ee8

Observation d174b87b-2a88-4461-8cc9-ae65dc60a418 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 2014

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.609130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:32:16.134721Z digest=sha256:72bfbbc08df0d1e9c849a949f711430c9ad970ed017e80740409f2ff148c50ee

Observation a80d550b-7804-4b87-a7ab-37bdeb6d35a5 · outbound

This paper cites On the Power of Multitask Representation Learning in Linear MDP.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning On the Power of Multitask Representation Learning in Linear MDP

Reference 3364

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.130092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.130092Z digest=sha256:446c0559761af58d0132da8a363bc1c76c6e77e6b4efc6ac0c2d0e60ebfb9d3d

Pith citing papers

Observation 1c0b5eee-179a-49b9-aa6d-0e3978524b3f · inbound

Phase Transition in Nonparametric Minimax Rates for Covariate Shifts on Approximate Manifolds cites this paper.

Phase Transition in Nonparametric Minimax Rates for Covariate Shifts on Approximate Manifolds Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:10.654724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:10.654724Z digest=sha256:a0c36b103ec21d8111bd716a98d9e8a4d9d78b7ca85b518b7f6abc0afdbaadc9

Observation 3f7af179-139a-430a-b8a4-f82a314dddee · inbound

One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL cites this paper.

One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-03T06:51:17.592951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:51:17.592951Z digest=sha256:ecc111a6c3aff0888c0d8e69ae35fb13104c7047416cb083c6ee3fabb57604cf

Observation d3d7ee4b-295b-4114-af5d-f0fd70ad5b99 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:27.343602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:f6a4e58b43cbe101696aa82d372de16418343f7983cf1f929b2094ff5276f1f9

Observation 84e593e0-db75-4a6b-94c9-5b7a2ec81661 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.999844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:26dabf80c62a3b2b2c5728f96be75a5e1c9a7aabaeaa02e081f3910f5266ebea

Observation 5a3779c5-27b3-4cce-88ee-6f1bf801dfce · inbound

Dual-Channel Tensor Neural Networks: Finite-Sample Theory and Conformal Structure Selection cites this paper.

Dual-Channel Tensor Neural Networks: Finite-Sample Theory and Conformal Structure Selection Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:06.993328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T07:20:43.847495Z digest=sha256:c7290580c86b77b6cf356471aba518f4ec900fe6df0c10a611583fe22c8c932a

Observation 15042302-ff5b-4033-9b07-7f2412602dbb · inbound

Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints cites this paper.

Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:08:12.160228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T10:07:03.605952Z digest=sha256:e75417dd3ac2b44d36dab46507e0a2f536b8dc2950325cb52023e5ac625b51de