Pith. sign in

Paper Citation Record · LEDGER

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2210.06718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.06718 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:45.725844Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:12.306741Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 29cab017-bd44-464d-b574-9ca054a70db0 · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.725844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.725844Z digest=sha256:f25392a3d2b300bb061c9e474e1228b8fde8dd4383cc0c71aae6f34551b92ab7

Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.439564Z digest=sha256:74e3b40a3343a57ec82fe350414ed79c5b346581491b3920f4b179e0bf74d9f1

Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.383535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.383535Z digest=sha256:998649d9defe1145ed866aa1cdaad690f9d66b69da559a3cb18b784bb3468eab

Observation 3e703aba-876f-44e6-abea-a1f2da435afd · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.259804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:5a9af1a18de23416843c8a0228774e30a533818e2e6b4a83de793d5acb7fd55c

Observation 767f32c3-af73-45db-9a21-063a76a9d3bb · inbound

Online Pre-Training for Offline-to-Online Reinforcement Learning cites this paper.

Online Pre-Training for Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:02.465774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:02.465774Z digest=sha256:1c7b67d21e87bc382cc5f15bdf91cb5496979d72313171b117556032dbb238b4

Observation 526b2c15-4c8c-4ea7-8dee-d7cd56523cfc · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:59.023990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:59.023990Z digest=sha256:56c0e0bb039d9b16450b7a6acf0a946b0cd35c2d5bbc81d47cbacf178d27eae2

Observation 54d71233-b41f-407d-9024-bd9fa7794d14 · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:30.243849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:30.243849Z digest=sha256:680e3a348a3f6ca85fffb6120572d6f8d3bd6b5776fa5d33741d318a9fd68eae

Observation 61191bb3-ed59-4c20-bad3-3192f7dd2288 · inbound

The Three Regimes of Offline-to-Online Reinforcement Learning cites this paper.

The Three Regimes of Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.709075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.709075Z digest=sha256:6bee81a5ed59c5335f49fdd644e1c795d12d3227ed7640dfb3890222f4695216

Observation f2dc4c0c-1335-4d2d-9833-b78d661542cd · inbound

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach cites this paper.

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:45:50.812731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:45:50.812731Z digest=sha256:217da45c24f901477ad7b2cbdcb021b6cd41cd4aaeb30e458fbe365bb0d0cc35

Observation bfaca071-6b06-4792-aa97-c858755680ae · inbound

On the Sample Complexity of Differentially Private Policy Optimization cites this paper.

On the Sample Complexity of Differentially Private Policy Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:45:54.791452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T04:43:58.646419Z digest=sha256:8464644f0ce31578c46eac94d8753788f5702ec75069beeb1f0a358803755123

Observation 5fe5774a-48df-4fd5-a21a-3bdccbd329b0 · inbound

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage cites this paper.

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:05:19.691967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:05:19.691967Z digest=sha256:f9449c6b2464a8361b9d9d7889e110bb2440fe020667b1a61c76c20c9be57cf5

Observation 0b234c6c-1aac-4e6c-8ffb-f70253b147a7 · inbound

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models cites this paper.

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:02.892690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T16:46:30.674244Z digest=sha256:f5ddb404d11b8e864acb0eaba0a7f325891a6a7e0ed628b292e84027ae5ae5ee

Observation 64a95906-4a9a-49a6-9fd2-4af53865b724 · inbound

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning cites this paper.

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.307546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:18:05.636131Z digest=sha256:5fe0c289fd3a2c9dfce0b8928ae5c62b579b8f4619ea784707a9e837232b329c

Observation 2de8b1eb-8296-4ad6-b7c6-3fdafcb22eda · inbound

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation cites this paper.

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:24.163245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:55:12.531822Z digest=sha256:23e679029237eb9e22aee28cbe6e4d9456b0a4272b6b89da48cba8aaa0375d3d

Observation 975e4bed-bd57-4f19-abf5-fb3ea33ed49f · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.640676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:5dc426f0c0070a6caa8440df7e5de42d136b266e3407e433063e0f87d8ab6df7

Observation 49d113e4-ba79-4d98-8959-46d1516a6769 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:12.446351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:bd829f7249e85537b814b8d4f4026530bf8b127840ceacfe472df1a37068d5e0

Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.342064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:0eb11da33242bd0ca4e8f2b3c2058d46fcfb28436c677bc3c82e253c4f32aef9

Observation 0b2b7717-b975-48fe-9716-aed5571caef9 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:08.508849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T14:54:30.895137Z digest=sha256:733a24d6e511a81207e550931c6621d68613f417cf6850df13f8e7b75a6d335d

Observation fdc32471-d64f-439d-bd74-0d8826c39e84 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:19:56.740644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T09:15:11.280343Z digest=sha256:5a2930d3a193bfbe6bd8e5ec3f9b4152364276b47a7efcc7b38d797b8df0c42f

Observation 796d031f-96a3-44c4-a061-5a62bb85114e · inbound

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift cites this paper.

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.988012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:41:20.147645Z digest=sha256:2f4a6041f02ac7b1d89107b9745db3bdb16a887a65972dbd9c36f7ffab971237

Observation fe4e1566-4459-4778-935c-a9e6bb1344db · inbound

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift cites this paper.

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:09:45.449771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:07:34.344964Z digest=sha256:721024321b9193695edd0bc5bb12bedcb70f671f8d3daf63e48fd2b37abf30ca

Observation fb0c0acb-c715-42a3-90ff-ba797ec27a00 · inbound

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization cites this paper.

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:33:27.192556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:32:42.972836Z digest=sha256:fdee18a7b0274d4f5e316408cab35cd7edf122d9d31d452b6074bbef81003201

Observation acb0340c-2a99-43bc-9711-85dba150d4d4 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.850150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:95ebb27a31d7ef818e40a1d3321e2dd00753780501cbc1392ad8c8c705d45aab

Observation 634974ff-59c9-40ec-a2f8-321d59b8e1f4 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.031118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:98c4be07dcbdeb8d237d547ff3c798119922e3b3ed7471aaa3811e794d81b080

Observation c40f1fc8-4888-48af-ab4b-b102ebdd2e46 · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.983617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:f8e191838bc5fe15e8e260fea71de4f24bf5ea92b35992147ef3b78aff63f29e

Observation 87b49c29-d6ad-4e76-8c37-c3c35c24bd8a · inbound

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation cites this paper.

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.308569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T19:29:00.117285Z digest=sha256:6221fc227cc340ff7a65c11c5af9eda870f1771251ea329a22aba000d6e36d65

Observation c84440eb-076d-4640-b8b7-a7d1d2dbf63f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 282

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:b90bf85c473cc7736fc2e003578a667f4e937929283e539feac6d59f40c46b87

Observation 8aa1f197-80c9-41b7-b6f9-9f1cb3891903 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 283

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:05.419064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:05.419064Z digest=sha256:5bf100a1ce1743ae9aab15a1e6579314bfcc6a7ffbcaf0af69a38eeb81d142a3

Observation 3ae795ad-3718-4759-94ea-ce72fd92563d · inbound

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics cites this paper.

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:13:35.351631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:13:35.351631Z digest=sha256:87508190c017958374b91a4d08c3f9dd942d39fb92adec0c60144790a595c7a7