Pith. sign in

Paper Citation Record · LEDGER

The Three Regimes of Offline-to-Online Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2510.01460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01460 v4

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:00:10.690877Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:15:07.735336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.874682Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ca2e4c2-6779-4f38-9a1b-28e2cd11ad15 · outbound

This paper cites write newline.

The Three Regimes of Offline-to-Online Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.167977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.167977Z digest=sha256:50288fda67bd81f5f5654619f54fc8ae87a43ebaf83375db3a6efd186301676f

Observation b4474f99-99c2-45de-a42b-e308f36073c0 · outbound

This paper cites Efficient online reinforcement learning with offline data.

The Three Regimes of Offline-to-Online Reinforcement Learning Efficient online reinforcement learning with offline data

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.268858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.268858Z digest=sha256:4183d35bf6c1d22f5575dbb659517d91582b0c8409481aba5f32e594025a5de5

Observation 8a974afb-4e9d-4839-b5a1-c3835a82ee5a · outbound

This paper cites Data quality in imitation learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Data quality in imitation learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.357293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.357293Z digest=sha256:44dc3dbaabe48f97e8169fa6cdbe8bb22444810a07eb65785be4ca7f8343e8ff

Observation 64f74df9-265f-46b9-b974-f7982ede9b8b · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.472452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.472452Z digest=sha256:b86de5e96abfdc72915c7c82e89a11f2890de8ef200cfa49ff2459be4435cf5a

Observation 88633603-02c1-4fb2-8bf1-05516eefb654 · outbound

This paper cites Loss of plasticity in deep continual learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Loss of plasticity in deep continual learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.588457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.588457Z digest=sha256:eb4d0f77e2be8ee0c81b3bda8dbe1c914ca1024d9097b5aa7bd863738d3baded

Observation 5456f005-3387-4bba-9222-69cef5bd08d8 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

The Three Regimes of Offline-to-Online Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.749438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.749438Z digest=sha256:7fe39625bdc288eeb70aa39784a8b670798c467993440f33be0f0245c703393b

Observation 44f3e1ee-45b7-44d9-bb9a-83644819cb5a · outbound

This paper cites Addressing function approximation error in actor-critic methods.

The Three Regimes of Offline-to-Online Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.916842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:06.916842Z digest=sha256:d2efb2bee114a7ef0f817455ba4e91b25ea8ec6f2fe37443ce20fa3bd6f01e91

Observation 7d6f22f0-1bc3-4bda-8214-587658d9c3cc · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

The Three Regimes of Offline-to-Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.065925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.065925Z digest=sha256:5bcfd6ebccedd2c572ccbbe4d1fe910080cfe2a2276741397bdb1bb0d275c46b

Observation daeae038-94a1-45bd-99c4-7e163bc04997 · outbound

This paper cites Bayesian design principles for offline-to-online reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Bayesian design principles for offline-to-online reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.185001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.185001Z digest=sha256:2bdd751153c6723369ac59a2e84c169245dd89e74ff63630fac4b86815bf21f4

Observation c19f7a39-8f59-442e-b3e0-843f87283869 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

The Three Regimes of Offline-to-Online Reinforcement Learning Overcoming catastrophic forgetting in neural networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.306912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.306912Z digest=sha256:cb78ee76016f428ed55d976514b183d4b7bcf658a5138cd7610c7f23533c1a00

Observation 79bd1747-8320-4789-9715-1e406d61c846 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.388006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.388006Z digest=sha256:b510276742790ac3a61c79e91ba72aeab8e8ea066d97ccf54647a9a75fe0331d

Observation fcbcfc6b-5b27-45be-8956-82845384997d · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

The Three Regimes of Offline-to-Online Reinforcement Learning Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.456957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.456957Z digest=sha256:a8bef8417921cbed090a265318fc20a4c1dcca6a2be2efd44ff4a3b54550cdf3

Observation 5fa3a69c-1c3a-49d6-abe8-1b2a4bd32e81 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

The Three Regimes of Offline-to-Online Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.505888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.505888Z digest=sha256:2d8d72d707f2d75040f26e72ef67b91deba1707fb9053647212c24e61741b02a

Observation 227a11f2-c231-4570-be1f-2d4df3c6b0d4 · outbound

This paper cites PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning.

The Three Regimes of Offline-to-Online Reinforcement Learning PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.558566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.558566Z digest=sha256:ff4af210c67152a6c35671898fa77025104daaf43472fa889d47efdc03f63cab

Observation 313ddf1d-bd00-4731-b934-2c4038123e26 · outbound

This paper cites Energy-guided diffusion sampling for offline-to-online reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Energy-guided diffusion sampling for offline-to-online reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.628640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.628640Z digest=sha256:06462036d5dccfd4d4275c7438afd0c6b91511cae099f4b1f5679c39ae696b50

Observation f94a57a2-8bdb-42a5-9f85-44c5b4c064e3 · outbound

This paper cites Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory.

The Three Regimes of Offline-to-Online Reinforcement Learning Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.642547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.642547Z digest=sha256:3b1cc83bfa8aad30b6f7317eb2a12b0f992ca47657aae31a1b0c20eb2009f0bf

Observation be610524-8dd9-4aac-8c6c-12f6ca48f3e0 · outbound

This paper cites The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects.

The Three Regimes of Offline-to-Online Reinforcement Learning The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.715050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.715050Z digest=sha256:b199d2eb6f1c13f87b1df785189b1dc60bb6d83ea7149735f97eab2940b9e6cc

Observation 39ffdfd2-a87e-4575-be96-2b4cc70eb862 · outbound

This paper cites Human-level control through deep reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.787156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.787156Z digest=sha256:1f4c4af749041ba1ef45c14ed18b4c8cc1d8bf13408351def2ca894a9156ccf6

Observation 4e7b98e4-c3d8-4ef8-ae37-279828667c32 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

The Three Regimes of Offline-to-Online Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.865376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.865376Z digest=sha256:768fa1d9dbf35223d6df9ef5beb5bb7657f9d3c33b8aae56a2ae8bc39d38be1e

Observation 31d5fcf0-dece-4f4c-a9e9-734ecd514799 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

The Three Regimes of Offline-to-Online Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:07.960413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:07.960413Z digest=sha256:08c65993265546bdeb88d7b3c8a4f6f4be6f2df439590c52bd5cd304003cdaaa

Observation 3a7bae6b-e680-4c8d-833e-94670839269d · outbound

This paper cites The primacy bias in deep reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning The primacy bias in deep reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.044669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.044669Z digest=sha256:a6daa27ab72730d4b576559565b5c57ccaf9765b77645313458513fcf715ce47

Observation 400e34b3-4842-4edf-9fd4-020fae7fdf4c · outbound

This paper cites An algorithmic perspective on imitation learning.

The Three Regimes of Offline-to-Online Reinforcement Learning An algorithmic perspective on imitation learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.086108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.086108Z digest=sha256:f18c2f7b5570f8fcb717ef679878cde88bf682faca1db6b3d4676e2dbcc8d4b1

Observation 77b01d49-ec9d-4562-8b06-0dd9590a8577 · outbound

This paper cites Experience replay for continual learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Experience replay for continual learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.092327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.092327Z digest=sha256:4511bbf77b87d0dcfdae8a77c35678a6b8682da736686670e13632508744447e

Observation 6d49d5e7-b6d6-4fac-a37b-93a914780dae · outbound

This paper cites Progressive Neural Networks.

The Three Regimes of Offline-to-Online Reinforcement Learning Progressive Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.162543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.162543Z digest=sha256:e064d650a268474443dc0cb2778a313492438edd4bac2b3cf33d81ea22508160

Observation f49e066d-df2a-47e3-8eaa-e10d4939dc0c · outbound

This paper cites Learning from demonstration.

The Three Regimes of Offline-to-Online Reinforcement Learning Learning from demonstration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.257104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.257104Z digest=sha256:0008ba817a8fb1f828a7efb6ea89c7ff50605479cbe81dafd301f56c980ef110

Observation 8baebfc2-c194-436e-b632-c7aab6fa00e4 · outbound

This paper cites Mastering the game of go without human knowledge.

The Three Regimes of Offline-to-Online Reinforcement Learning Mastering the game of go without human knowledge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.382285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.382285Z digest=sha256:64774c9cded3bee8cf4d020cf703fee87ec9e35f3b27387e5ded58cbf4664a80

Observation 746071d9-dffe-433a-9a4e-3d668e546014 · outbound

This paper cites The dormant neuron phenomenon in deep reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning The dormant neuron phenomenon in deep reinforcement learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.538793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.538793Z digest=sha256:d0ffcfc58ac18c50b353d644f4898d57e5c509689c942e4a62f454a126f4fabc

Observation 61191bb3-ed59-4c20-bad3-3192f7dd2288 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

The Three Regimes of Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.709075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.709075Z digest=sha256:6bee81a5ed59c5335f49fdd644e1c795d12d3227ed7640dfb3890222f4695216

Observation fce4ee65-e198-4c01-94b6-eaafdcf147e3 · outbound

This paper cites Feedback in Imitation Learning: The Three Regimes of Covariate Shift.

The Three Regimes of Offline-to-Online Reinforcement Learning Feedback in Imitation Learning: The Three Regimes of Covariate Shift

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.940601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.940601Z digest=sha256:2c090e3297eb3a9696ac2497d726762e6ccd4d09040f6fcf4464c8d2756e5211

Observation 506f1782-cc61-4490-ae3c-baa7ff54259b · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Revisiting the minimalist approach to offline reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.188443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.188443Z digest=sha256:1bbe4065436fd0d1ec4db4e4cd052de41bf4f6abcc953a351cccccf02ad38c52

Observation 78036c47-8bf7-4ad0-9557-ded2448a01dc · outbound

This paper cites Jump-start reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Jump-start reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.366372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.366372Z digest=sha256:2b5f629af2b3f4d228afc2991b9fd1ce465b39f01f6aac99005cb78ab1acbb18

Observation 5f03b118-fe6d-4584-bc95-413a32e2d38a · outbound

This paper cites Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.551445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.551445Z digest=sha256:b19443b6dce6c213331ebf0301f171d7486b0810e16b429a62bc9d81f0b2306d

Observation dcb8cee3-b4ea-4d9f-a664-898332447b5b · outbound

This paper cites Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem.

The Three Regimes of Offline-to-Online Reinforcement Learning Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.798532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.798532Z digest=sha256:8f07605d2d1cb9b6e077d0d0ddf7c80b69414ff5c36c02720d11cb0391f18664

Observation f83bd745-f59f-48bc-8cb9-c5b98b025e48 · outbound

This paper cites Policy expansion for bridging offline-to-online reinforcement learning.

The Three Regimes of Offline-to-Online Reinforcement Learning Policy expansion for bridging offline-to-online reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.968342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.968342Z digest=sha256:e77aed7eb75a996125686b4b0f9ec3f5fd48a72f733fe10eefda30f8d5cc477b

Observation 960fc39a-6c78-45a9-8cd2-88d8b797036e · outbound

This paper cites Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data.

The Three Regimes of Offline-to-Online Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.173421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.173421Z digest=sha256:ab2daee4720a65dc4d944f533d47aa4689e28ca8dbfe84453e5d4dde2f109776

Observation bc4c0827-f1c4-4e30-ab56-44fee2f15ec0 · outbound

This paper cites @esa (Ref.

The Three Regimes of Offline-to-Online Reinforcement Learning @esa (Ref

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.349681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.349681Z digest=sha256:5e162205e7ad732eed5e7527420453b9743bdd8e1a4a131a6e28d45428656ae6

Observation e0c5250f-5091-4b8a-8cb0-da5f49b64545 · outbound

This paper cites an unresolved cited work.

The Three Regimes of Offline-to-Online Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.517374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.517374Z digest=sha256:267bd6dfd217246d1e1b514599cf6f87deefadc2ae835201280edcfdb58378f0

Observation d42a8413-8f1b-4bd8-b03b-71a19c49ac2d · outbound

This paper cites an unresolved cited work.

The Three Regimes of Offline-to-Online Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.690877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.690877Z digest=sha256:977d4822d91effeb0ce6c2150b04429280013073f82589fbcf3bcc40ef4f2967

Pith citing papers

Observation afa959f4-104c-47cf-becb-304ceca00e76 · inbound

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization cites this paper.

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization The Three Regimes of Offline-to-Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:12.544546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:32:42.972836Z digest=sha256:fb5c518cbc1131bc03bc2ce68164a2c748b6916a37be544173e2bd9278456298

Observation c25873e5-3c82-4fe2-b2c1-a573384893d3 · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors The Three Regimes of Offline-to-Online Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:12.544546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:a92897aa4bd1432328e24f7d9642896bb695c8b559aff05d4b58751f898ab2a9