Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning with Offline Data

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2302.02948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02948 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.188595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.181580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07b8cd2f-ab37-4319-9b16-b246748fc2b7 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.436566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:5f233b1640c1db8e94acce5f7c95442678150c86343d663d32c2f7b1e33b568d

Observation da1fdad5-7c8d-4786-9ebf-e5c3fdaf1061 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own Efficient Online Reinforcement Learning with Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:44:02.559846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:b80ff5b8d4dd1660589f0923cec2058d4389a0b70f0cacae2e792611c1e17681

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:a255c3399e050af2f98f41e0b4c26cd751c9ea505369a2c05fe136810d780175

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:98514d551e241c53c2bbbc261adb826aed9aefbc7baadce3131adb6862a39438

Observation 19b8f177-1080-4532-812f-1a5351fb1d6d · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:28.644931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:28.644931Z digest=sha256:ef558bb04a734d2cc5bacbdc529fb411087e06c0b7dead9a62dcb3a3320272c0

Observation 6af8edd1-46b1-4419-882f-c5d3db83d8a5 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning with Offline Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:26.920260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:26.920260Z digest=sha256:9c252cfe2e74fc1a06f29dca895be154ea583244496c9cc6c9886d2b1f207c6c

Observation 8ea21e37-3b0e-49ab-9fd5-8a67e61b2199 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.613254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:04:41.766964Z digest=sha256:9497e6daa95409d3801ddb4a5a5c5ba003765531512417f771d771a531dce111

Observation 7a7d4669-fe08-4290-ab25-458ba3ca73e6 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T21:40:44.030470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:40:44.030470Z digest=sha256:2ba4a0d01c65a76b013ff44210cc48dc614e27e8b17c3c67c328f63258ec39be

Observation 93d52c32-20b8-454a-bc34-a47bfcea627a · inbound

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation cites this paper.

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:06:52.898450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:06:52.898450Z digest=sha256:692d3d910440b71d6bb13823869dc6d1a3fdbb89116e0bd863217130fee2beee

Observation 44b8beff-9de2-4cae-95ce-d09ce9fe70b4 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.108166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:b9b498f10c73fcbc37a10c453275cbcb6f6a26b79c591f6b842b67561103a4f8

Observation f2bb1677-e9ad-4059-94aa-336f814d5aae · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.704833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:716f1f6dcfeef9dc3725226bb42c879c4acff998b7c459dee3a61e751580c328

Observation e4ef87bb-5bf0-42bd-9790-d38f19d2060e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.164807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:dfdd72cd61c0f8537ad85174e3ce0100716b1a53e6063ef461750fcd90c5566e

Observation 4302685f-03af-401c-bfd5-d9e21e29b59e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.785098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:de905bcb37bdb9952d4087f5705241c0be54050d2c6a7ab1d5d0bfda9f48246f

Observation d19e3344-8788-4200-ae8b-6b7fb000eb31 · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Efficient Online Reinforcement Learning with Offline Data

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.803777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:22e1eda37f8e564528aa52d21d62e11ea7c05dbdd4d9fc0431ac6125771e0cb0

Observation 3122a81d-378e-485c-97e4-c1ebfff421a5 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Efficient Online Reinforcement Learning with Offline Data

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.514884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:97ef5fcb9b00256d4289898127f93af556087db393c668fd30ba17fa821234e1

Observation 3bf457a2-ff38-4369-8ac5-5166c4c1a1ca · inbound

An Introduction to Causal Reinforcement Learning cites this paper.

An Introduction to Causal Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.183066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:17:57.091481Z digest=sha256:5e569806e42691925e8dab9f3a7afa457faf1f19f46f455fd5ec9848eb3d05f9

Observation 63993894-2958-492b-a60f-a338fef6ccec · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning with Offline Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:eb7b1cfa2feefacf2401c3cf702a948207b3c86a8386a2a4c58b8f21555655c5

Observation b0587974-934a-47ec-a15b-e711c698d88d · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:21.826493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:21.826493Z digest=sha256:d4f2593c7d1bef0cee678d3edf9f98544cf9999833d8953ac2ca72a0b60c1ccc

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:a991eaf0d8833636ded48efd467d1fe999c4f4c03cd1302d8163e1b967e86f62