Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning with Offline Data

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2302.02948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02948 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.188595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.181580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07b8cd2f-ab37-4319-9b16-b246748fc2b7 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.436566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:d4e8d5b675f42fce2bffb92da9d01bfbcd738c489e638ca06797bf999938ab43

Observation da1fdad5-7c8d-4786-9ebf-e5c3fdaf1061 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own Efficient Online Reinforcement Learning with Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:44:02.559846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:7f476731a35f42e6e2a0d0b6b6f97beefc70fe571376f893e1489a8845036508

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:a255c3399e050af2f98f41e0b4c26cd751c9ea505369a2c05fe136810d780175

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:98514d551e241c53c2bbbc261adb826aed9aefbc7baadce3131adb6862a39438

Observation 19b8f177-1080-4532-812f-1a5351fb1d6d · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:28.644931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:28.644931Z digest=sha256:ef558bb04a734d2cc5bacbdc529fb411087e06c0b7dead9a62dcb3a3320272c0

Observation 6af8edd1-46b1-4419-882f-c5d3db83d8a5 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning with Offline Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:26.920260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:26.920260Z digest=sha256:9c252cfe2e74fc1a06f29dca895be154ea583244496c9cc6c9886d2b1f207c6c

Observation 8ea21e37-3b0e-49ab-9fd5-8a67e61b2199 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.613254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:04:41.766964Z digest=sha256:235b9e6baab7b428ed23a363001340b697a399d480aae30869dac5971f260ee0

Observation 7a7d4669-fe08-4290-ab25-458ba3ca73e6 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T21:40:44.030470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:40:44.030470Z digest=sha256:2ba4a0d01c65a76b013ff44210cc48dc614e27e8b17c3c67c328f63258ec39be

Observation 93d52c32-20b8-454a-bc34-a47bfcea627a · inbound

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation cites this paper.

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:06:52.898450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:06:52.898450Z digest=sha256:692d3d910440b71d6bb13823869dc6d1a3fdbb89116e0bd863217130fee2beee

Observation 44b8beff-9de2-4cae-95ce-d09ce9fe70b4 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.108166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:88a7776a443230c01473c6acf23aec65961ae8e6658db5889c6c21975cfb9659

Observation f2bb1677-e9ad-4059-94aa-336f814d5aae · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.704833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:37cc026010591e6e333e3e23e7702e7959857ac336cda38c5cbf22f5d347622c

Observation e4ef87bb-5bf0-42bd-9790-d38f19d2060e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.164807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:3a8a06b39633d5b71320240419fcfeef75995fb7785b020b73d37d1e02f5a1bc

Observation 4302685f-03af-401c-bfd5-d9e21e29b59e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.785098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:47e93fef7acb5e5372eebe4bb8a16d4ab8c9fd9d6a4bf2848d27fc734d7e4919

Observation d19e3344-8788-4200-ae8b-6b7fb000eb31 · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Efficient Online Reinforcement Learning with Offline Data

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.803777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:50bc9df69c331ee233820a1845c9f938674c322a13b7756afdf3670b95b69227

Observation 3122a81d-378e-485c-97e4-c1ebfff421a5 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Efficient Online Reinforcement Learning with Offline Data

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.514884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:b62fc13299936057e8c92cebf09944e134d37acbac3bdbb8afd5a90fcad446e4

Observation 3bf457a2-ff38-4369-8ac5-5166c4c1a1ca · inbound

An Introduction to Causal Reinforcement Learning cites this paper.

An Introduction to Causal Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.183066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:17:57.091481Z digest=sha256:de4630baff8e2f239d7edfd7ada61b4b001f18db2a882c252989056c7d0ab274

Observation 63993894-2958-492b-a60f-a338fef6ccec · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning with Offline Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:eb7b1cfa2feefacf2401c3cf702a948207b3c86a8386a2a4c58b8f21555655c5

Observation b0587974-934a-47ec-a15b-e711c698d88d · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:21.826493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:21.826493Z digest=sha256:d4f2593c7d1bef0cee678d3edf9f98544cf9999833d8953ac2ca72a0b60c1ccc

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:a991eaf0d8833636ded48efd467d1fe999c4f4c03cd1302d8163e1b967e86f62