Pith. sign in

Paper Citation Record · LEDGER

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

As of 21 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.13768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13768 v3

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:29.861150Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:11:16.568415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:13:18.015262Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c895a0b-76a4-440d-8973-a3d0e758408f · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.242665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.242665Z digest=sha256:e9102b5624eca766472209de2fa419c687b7f42bc111808f53755bfe2f3044bf

Observation c98ec8b3-3062-4b4e-8bd9-9074a259a4ab · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Analysis of thompson sampling for the multi-armed bandit problem

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.249582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.249582Z digest=sha256:01a4f5e6bc9445ad21ba700befcfd2bfa6464ea9ce68eac1c1b8dbc9bb1d814d

Observation dbd2d4cd-c320-46f3-a782-489dfb43dccf · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Thompson sampling for contextual bandits with linear payoffs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.255372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.255372Z digest=sha256:88ae02033853dbdce5f2dc2d94094bbcd7738b4a12d2bf692de68b446b9a9d82

Observation e1ed65f2-ecaa-4a58-9c66-3dbd4c185620 · outbound

This paper cites Optimal Best-Arm Identification in Bandits with Access to Offline Data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal Best-Arm Identification in Bandits with Access to Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.266065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.266065Z digest=sha256:c632c9e14dd6ad3c6f8ac92617e67e81f478dc02d4e17adb2743fdb201e2697c

Observation fb1ea420-bc2c-40a9-8b9f-0a2a64753cc7 · outbound

This paper cites Exploration--exploitation tradeoff using variance estimates in multi-armed bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Exploration--exploitation tradeoff using variance estimates in multi-armed bandits

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.274118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.274118Z digest=sha256:625721385845ca8f0620bb4c7690f7b06092528b7cea3eb8f223f03d360a6818

Observation 01e262d1-166e-40a0-9384-e84988a08afc · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:31.817957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.279590Z digest=sha256:756c3d597b7930195b975c5fc3669c929b90fb0013a53121020a34f9e531e399

Observation b9a12319-d135-458e-9167-e33af680928d · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax regret bounds for reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.287028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.287028Z digest=sha256:bce5c575b774ec4926afe87c2b3b3bfdfcc6885878dbd24e48c7454bc8c6d153

Observation f44891fb-e7aa-4523-b982-6c7d55e07e19 · outbound

This paper cites Stochastic linear bandits robust to adversarial attacks.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear bandits robust to adversarial attacks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.297657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.297657Z digest=sha256:91da5513109cd1e3b53216310fd8ff1f7b6c1b22ed22ae725cf3b831ceaeaa9f

Observation f2bb5257-693c-4215-80da-4b2166df8393 · outbound

This paper cites Offline contextual bandits with overparameterized models.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline contextual bandits with overparameterized models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.757168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.307453Z digest=sha256:3fe172fd54771d477c1603af7351aa61451f1920e3d4ce35c68fa9b77e27054f

Observation 6f739da6-b875-4918-a870-78e524e0581b · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.321056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.321056Z digest=sha256:1567549d5842f902a7fe517b138245ff383fc2933ba142ddea9e1179eb01aaca

Observation 92622cb2-85f6-4433-b05d-c2d61d0a5f07 · outbound

This paper cites Kullback-leibler upper confidence bounds for optimal sequential allocation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Kullback-leibler upper confidence bounds for optimal sequential allocation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.715206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.327688Z digest=sha256:72274ec30f141c204076f1231b87ab48ec30573dfeba5a3d0e74c9cba9f23818

Observation a0f51a63-b927-4230-a41c-a09f73794017 · outbound

This paper cites The Elliptical Potential Lemma Revisited.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The Elliptical Potential Lemma Revisited

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.334478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.334478Z digest=sha256:fbbaafa224e2e81894eb84803aeee21319c0ecd89cdf4ae28e211c2c7a3e8adc

Observation e6ef10ea-4654-4bf6-880f-6913f90944ba · outbound

This paper cites An empirical evaluation of thompson sampling.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis An empirical evaluation of thompson sampling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.343000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.343000Z digest=sha256:03370c3f121ebd1a9e893b8aa608ca7d00e177c03128707a5266f38e794fd255

Observation 49cc509e-74a7-455c-adfd-0ed56063cc20 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Information-theoretic considerations in batch reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.350532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.350532Z digest=sha256:9fcf701938eb116553138c12b24a255505bff0ff6f6d0e3b17f89e957a18945f

Observation c42f4de1-f0de-40c9-89d5-596efbef1efa · outbound

This paper cites Leveraging (biased) information: Multi-armed bandits with offline data.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging (biased) information: Multi-armed bandits with offline data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.359279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.359279Z digest=sha256:6ef8522d93b1abbd925ba6ff1bef3455609a23b27517638ea76bf484c12e0a8b

Observation 0d4484bb-ebcd-47bf-ac44-dc2c28e5298d · outbound

This paper cites Contextual bandits with linear payoff functions.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Contextual bandits with linear payoff functions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.367609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.367609Z digest=sha256:8da7e5f0245e86bb0cb349e5500e3f2d2c465d81086ce84c7247d874eb472d29

Observation 4f273f8b-d01d-4bb3-8e7c-f73427d23de8 · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Stochastic linear optimization under bandit feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.623429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.375257Z digest=sha256:0190746d6df080ef595c8b9400bedecdb74cd3336965d53486d5ade04befc44c

Observation 04905289-db52-4afe-b818-25bbb45f9a4a · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax-optimal off-policy evaluation with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.595939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.381472Z digest=sha256:f7f59cb02025a2b26320b98dc0bc2592b060a960520fa17f6a1f50f76f0331af

Observation 010d2342-131d-49e0-8545-81ebdaab5c71 · outbound

This paper cites The kl-ucb algorithm for bounded stochastic bandits and beyond.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The kl-ucb algorithm for bounded stochastic bandits and beyond

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.570089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.388370Z digest=sha256:b99fcf747f86265a0f0bf2350134e6b049db91e30ea899ed562430ee15599746

Observation 1eda491f-0c24-472c-ac81-c687a01adbb1 · outbound

This paper cites Guidelines for reinforcement learning in healthcare.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Guidelines for reinforcement learning in healthcare

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.534809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.396228Z digest=sha256:67231d9d05f472146eea6ce77169fd5108342d59f1dea6c26f0fa56f70ab9222

Observation 5ea1c3bb-61bf-45f5-b3c4-c629bc3c2ccc · outbound

This paper cites The movielens datasets: History and context.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis The movielens datasets: History and context

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.404480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.404480Z digest=sha256:c5f0e0e23cd443e94b9c3b761da3b8e4e713e51e8e47d50bd16b2d144fd2ce2d

Observation b265bfe8-4724-46cf-9f59-ffb10609395e · outbound

This paper cites A reduction from linear contextual bandits lower bounds to estimations lower bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A reduction from linear contextual bandits lower bounds to estimations lower bounds

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.483752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.416232Z digest=sha256:ea25a18b76a971f5a2df546c9477cfa69c627d89e64f2be452d8f7b3af934efd

Observation fe546f33-8415-4f92-993b-e61d5dfc2339 · outbound

This paper cites Deep q-learning from demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Deep q-learning from demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.425884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.425884Z digest=sha256:aca52fb6c73d8464c50ecd1bd1b83491cc7872c790e3435078ae460134e2aa4c

Observation 4d80e72d-04bd-4be8-ab07-65ff2124aa24 · outbound

This paper cites Optimal best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Optimal best-arm identification in linear bandits

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.433662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.433662Z digest=sha256:c1cacdaa3031dbdec5ede2b632ee0c4606ccf315ee2acee45da6f01d7030f5d1

Observation 582dc0a1-c169-4ace-9248-d63eb78cc5ae · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provably efficient reinforcement learning with linear function approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.440235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.440235Z digest=sha256:e83b31316688dd6f7ca044bce823f5ff52594c60d5baeb5992c2b7cda0d72141

Observation 549a36c1-cb33-402e-a0db-ded1c5cc646a · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.447002Z digest=sha256:d01eaccbce4d5e6018a738671f1720fc91ec97f88164f6bc69f14add0324f6cb

Observation eb141d2d-c21a-4f56-a2aa-89f723036659 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084--5096

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.453644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.453644Z digest=sha256:4a0dc03726e22156c98082fbe652c1e9f28b29b63f0dc3bf97308511d6c58647

Observation e7ac0e9b-76de-4530-9ae9-4ef74f1c9216 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Conservative q-learning for offline reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.460288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.460288Z digest=sha256:3034e80f29015046fbe9609fb8ff725ba08d273f84a7b21e37fb13f615ea6308

Observation 9dd65cf7-1b34-4045-80ca-b603d1c277c5 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Asymptotically efficient adaptive allocation rules

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.323438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.467738Z digest=sha256:7dde9c87a6d8555b5ea284d479c3b6f5552686d305cab0e288638bbdcf1db668

Observation 48770e6b-2a6e-4cbd-a145-19f174a6274f · outbound

This paper cites Bandit algorithms.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandit algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.476237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.476237Z digest=sha256:d373e9334bed7a183ed4f17574d325c086b499805d5558f07517132896c7bd4c

Observation 92c2d990-9539-41b7-8a02-2a9f06ad463e · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.483421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.483421Z digest=sha256:6b0af7e1e7756318232d62f7d06240e915e0603f27067b7ac85b298cba6b7643

Observation 1fd1f3f3-1d2e-4905-9663-0cde19d7c110 · outbound

This paper cites Lee, Yuejie Chi, and Yuxin Chen.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Lee, Yuejie Chi, and Yuxin Chen

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.258080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.490173Z digest=sha256:d02e7eca506f0ef970878fc700e535b05703fb4042ee05d027ad42d64944df8a

Observation 44abf8ef-5c8d-4d7b-bf1a-d4459351a26a · outbound

This paper cites Settling the sample complexity of model-based offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Settling the sample complexity of model-based offline reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.229101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.496105Z digest=sha256:a3d8484494623ac5ab3f4b3775aeebac767955c3a27aab019fc82ca2670bed18

Observation 6b53940f-e37c-4738-a6b0-c4f072530764 · outbound

This paper cites Pessimism for offline linear contextual bandits using l _p confidence sets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimism for offline linear contextual bandits using l _p confidence sets

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.208203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.504425Z digest=sha256:58fb74652c34c6c37f335f8236b46b094d039560f35d6bf9aa6f3f7d6ec1aa53

Observation e61f07a1-673f-495b-88fb-876bcfaadb62 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A contextual-bandit approach to personalized news article recommendation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.512902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.512902Z digest=sha256:50fd6191a5dfedb036b764645e5c4a3b60e1f70e5c69313779a4f7bdae8928a6

Observation 987369e1-50d9-462b-ab7c-2165fc6a5f6b · outbound

This paper cites Fast active learning for pure exploration in reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Fast active learning for pure exploration in reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.159599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.520932Z digest=sha256:c1d16411ecb91cc1ffe54154fa0f1e4df6e9cc2401a2f99ae97bf8b5a96ea11f

Observation f557d306-35f1-454a-b0b5-d501c09e3a6d · outbound

This paper cites Efficient memory-based learning for robot control.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Efficient memory-based learning for robot control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.530093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.530093Z digest=sha256:475700160680cc4d4de28b8721975d2a3b431d86643afb3d17fc0a3bfe079616

Observation 548bff10-8852-4329-83df-625b62a4ff98 · outbound

This paper cites Collaborative-filtering.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Collaborative-filtering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.105132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.539479Z digest=sha256:ef582b0516488a29edb6ad9223579716062883741aea42ec4db51501c89a2a4b

Observation c9ac7849-878d-47fc-a732-d1156fe9ac2f · outbound

This paper cites Finite-time bounds for fitted value iteration.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Finite-time bounds for fitted value iteration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.547938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.547938Z digest=sha256:12f2bcdb4c9a82fe70cc5d4f7b385986af7c84de062be82ca111899082423878

Observation 70166982-94ee-4140-930b-0fbe349720a1 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Overcoming exploration in reinforcement learning with demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.556309Z digest=sha256:8f6b3b416d0737fb836d3767099a4c626c581f3092fb3d395b30497a0abef6cb

Observation f668b8b1-7887-44cc-91b5-c1c598401423 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.565639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.565639Z digest=sha256:62b00c2a1abe4d42de0ce9035d725b592f87aae3a6a69458944fa70fe3e4126e

Observation a68c2aa5-0991-45f9-96ce-9013e9e23a41 · outbound

This paper cites Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.572118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.572118Z digest=sha256:0bf589ba59a94d4ee101e283e58b40b0c9aac1a42e8fe1abd4a2fc4903d6a9d6

Observation 42d21d93-8346-44dc-a5be-0bd20480f589 · outbound

This paper cites Cutting to the chase with warm-start contextual bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Cutting to the chase with warm-start contextual bandits

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:31.008409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.579950Z digest=sha256:b2b59f4ee1dc155c2166820eb13707f28a44c23331b9df29d8c9cb779c2e8b0d

Observation a23cd302-2197-4e6e-a6d0-ab54448f3761 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.587974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.587974Z digest=sha256:509a3b8456ac88728efaf5148892a7338f778862bbd323cbcd09af5ee945b96c

Observation c2f7c428-45d9-4047-8262-7d8052ce2572 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.981176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.594940Z digest=sha256:8078de666cb0da3ee4edf2c973f5b0e1b4b30cdd832a7dab34a822569b9a5bfc

Observation 6c010a11-7c86-4337-8617-ef321ed93183 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Agnostic System Identification for Model-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.600345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.600345Z digest=sha256:55fdffc1f3c4353d60792961f20b8d025e11867b4225ea29d84039d4dc212a85

Observation ef0b3f80-f1e1-4e99-989f-385b3fdab78b · outbound

This paper cites Bandits with Mean Bounds.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bandits with Mean Bounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.612927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.612927Z digest=sha256:44c2a635408b985231dab4eae449bd99be55203f8e982755c086729d3cb42850

Observation ec621217-6fcc-48d3-98c6-c099f076e4ee · outbound

This paper cites Multi-armed bandit problems with history.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Multi-armed bandit problems with history

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.946397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.622678Z digest=sha256:10f6a2ab89796d6a8b5b30de9d6165ed97975c3f1f83244e33861703539ab7d5

Observation 796f590c-eb96-45de-829c-327756193768 · outbound

This paper cites User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis User cold-start problem in multi-armed bandits: When the first recommendations guide the user’s experience

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.632051Z digest=sha256:c45381d12488f647ce393374c746fa02a70ec5141a05c456a067ed93fb057579

Observation 072ccc22-bb4b-4321-8cbd-af8035632d12 · outbound

This paper cites Best-arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Best-arm identification in linear bandits

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.869269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.648294Z digest=sha256:1bb4260bf09a14d705d8e01d330b3f227635fa4de3cd19a7ad55a976306e52eb

Observation 905dd96b-4a67-4054-979a-a5edb08b2da2 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.656847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.656847Z digest=sha256:1a3c0794180f198298c630fc15297f04f469a8705a6d30760fbe18eefa601cb9

Observation d557b812-f358-4906-9f25-cbfe7e9a1778 · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:30.840977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.669284Z digest=sha256:bbb030ab79a1a39394d024b666a8079f8215c2ddc7d2c2f6b45a209ab408efcb

Observation ef0ef3db-819e-4f9f-995d-2cdc6b32b605 · outbound

This paper cites Algorithms for reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Algorithms for reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.677081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.677081Z digest=sha256:b33e41c1c64f7b8e592a815dc305ae1fc3370eba37f33ad8298889eab1713caa

Observation 4fabacf2-bb9e-4fca-b173-82e627b3c8fc · outbound

This paper cites A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.160097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.685128Z digest=sha256:9a288e841dfbd943566363d6f841319b9fd03e16fa9c0d52daba5b863cebe486

Observation 55fb2b55-bc45-4ca8-842e-fe744e94b712 · outbound

This paper cites Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.691317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.691317Z digest=sha256:6e3b3b33d394e9b952716e5ee7214c8dbc2e57022f6385c435b7b726737c1b01

Observation df602222-d459-41ae-9e10-d48e203e50b1 · outbound

This paper cites Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.799656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.697435Z digest=sha256:395bad4d59696f3c23cc7606cfb5038f029dfd3dba0cc1b3e193f9ecf266bb03

Observation b0332b30-f19e-40ce-8a94-878cb945dace · outbound

This paper cites On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.705219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.705219Z digest=sha256:e3792cd74b278cb47d8dbe42ae8a024108a5c65362a9bad0cb416b54745da738

Observation f3f4b2b8-416b-4af7-97f8-942f265df193 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.713695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.713695Z digest=sha256:56a2b691b0a40c2dbc99299640b005d3a9b690d7b0968e3beadd2baef2cbbef5

Observation a9c0dbe8-6e5e-4322-9442-971d1fa2aae3 · outbound

This paper cites Pessimistic model-based offline reinforcement learning under partial coverage.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Pessimistic model-based offline reinforcement learning under partial coverage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.753059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.722442Z digest=sha256:9300661577a32a41c343a23ef30f579bb19dca44e98878cf71cb46711d5ca922

Observation 3cac87ee-b04c-4e77-b05d-d5f45757dcd9 · outbound

This paper cites Representation Learning for Online and Offline RL in Low-rank MDPs.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Representation Learning for Online and Offline RL in Low-rank MDPs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.732840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.732840Z digest=sha256:69afcfcea297c69dcea57bf4d4442ca9df1f879ad3197fc65d9799ef65293d2f

Observation 87db7217-299f-4cca-b768-227f46ae0dfb · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.744641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.744641Z digest=sha256:cc74b3d47345db1bf21f6873857814e87339e95e659f1087f8ca6ce574657a58

Observation f508c3f2-438e-48e1-b5da-9d5d4cee3c12 · outbound

This paper cites Leveraging offline data in online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Leveraging offline data in online reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.710969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.750323Z digest=sha256:843613dafa67beea899ddc51c4280c04fca7ace0098ce8770a399645d3edb372

Observation 685e482b-1df0-4cd1-9d5c-2a4c6f90efbd · outbound

This paper cites Experimental design for regret minimization in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Experimental design for regret minimization in linear bandits

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.689024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.758386Z digest=sha256:b54be323fa30cd2101b9699a714642dae1fb1ffa53c16e0244581455b1037e9c

Observation 09163372-f577-437b-8cb6-8605aac2909b · outbound

This paper cites Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Oracle-Efficient Pessimism: Offline Policy Optimization in Contextual Bandits

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:30.031698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.768398Z digest=sha256:2d622582c55537fcfbd2be2c58cc3126cef4571814ff52047c25f01d872058d6

Observation 0716a315-be0d-4d68-8c23-e9ece9cd99ac · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Bellman-consistent pessimism for offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.665695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.778054Z digest=sha256:cee4efa5a11b02482a80ea169afb70a584a9aa3264cf7676a6fd87e05595a1ae

Observation d779f11e-5fb4-4a27-b7e3-099c6753d832 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.642102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.787656Z digest=sha256:048bd79383820e0eaafc308f5d82edc1f5c8863cb75045f27480efb86b95fe25

Observation ec161056-9f73-4450-8caf-ccb1b88a4a80 · outbound

This paper cites Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.794976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.794976Z digest=sha256:f16012356227169e994b5a6c864e67839e460ec2507b08d37c7d0e5fe4d349b8

Observation e7181073-99bf-43f7-83f0-2b3174dbdbba · outbound

This paper cites Minimax optimal fixed-budget best arm identification in linear bandits.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Minimax optimal fixed-budget best arm identification in linear bandits

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.801662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.801662Z digest=sha256:9d4810dffcfde897761d7940dd33c691bb87cd54769badf57cccbbc9077ca04b

Observation 5d7e2902-3ce0-4eec-bcd5-fe73934058d7 · outbound

This paper cites Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Offline Reinforcement Learning for Wireless Network Optimization with Mixture Datasets

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:21:29.968928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.810991Z digest=sha256:f2e0e5a29f5bb7e6dbcefbe99d29c88582562dd5d56705ce8fec2ade033abecc

Observation ff14bda5-a614-463d-96cc-aba9a321377b · outbound

This paper cites Provable benefits of actor-critic methods for offline reinforcement learning.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Provable benefits of actor-critic methods for offline reinforcement learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.605103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.819313Z digest=sha256:5a3c854f5d0cd30a2d3522b5f4983d117b3d1f1110d05d8aecbda844e0657c38

Observation 4af4ccfb-fb8b-4844-9c47-dd6400cd56b6 · outbound

This paper cites Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.829348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.829348Z digest=sha256:b5426eefed6f47b03f017cc98342a27a228cb5ccb0b5e2fa12a3b3eb9d59d1a9

Observation c67f8696-c5ea-46ea-ab54-995e8465a604 · outbound

This paper cites @esa (Ref.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis @esa (Ref

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.840130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.840130Z digest=sha256:62cf4a6f1fc555487a2a96a7bd4bd64089c854ec2f9a3a813fc09cfe0ccfb877

Observation cdcc5227-6b28-4a7b-9012-800c7f84a40a · outbound

This paper cites an unresolved cited work.

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:29.850447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:29.850447Z digest=sha256:6997a193bf5e2dbbe7e3bc511d6f12a00f80e2a185fc30d80816633b7e3d345a

Observation 9079adcb-e6d0-4853-b3f4-ae45c8c1c524 · outbound

This paper cites page @startpage numbered @text Submitted to @long ( @short).

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis page @startpage numbered @text Submitted to @long ( @short)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:30.535409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:21:29.861150Z digest=sha256:66fd75d8f914fde63ea2badd56fd99bc9abe2a0515e95c45fca3f4865b6d93da

Pith citing papers

Observation c593b69d-45c1-400a-8621-a3891331fde4 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.017226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:ff5aa2cee3a1b30cca8ad670408d3f3b8197680fc65f524f5d58e47b4847ba75