Pith. sign in

Paper Citation Record · LEDGER

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.16856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16856 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:46.595010Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:11:16.568415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:13:18.063828Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7312754-442e-45ba-a88b-ee4ac8fec3f6 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Is Conditional Generative Modeling all you need for Decision-Making?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.563327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.563327Z digest=sha256:5fc0c26862af603ec122938b979f1646bac055bf1a050452c526d81049f443a1

Observation ddec04fc-0c8f-4096-b3bf-0f0d762108ec · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.985644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.985644Z digest=sha256:23986f1b7c442ac24f02e478d6c01fee6ecd5dc8f2662519e6c8dca95affc4ba

Observation fdffa202-d7d7-4237-889c-d54e5cf02b8d · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Offline Reinforcement Learning with Implicit Q-Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.287992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.287992Z digest=sha256:99394e85fa7d2b9076d098150b9eb9f5cef691be164dbbe8a5a3d1a3b14e848b

Observation 1227962e-e897-4bac-8abd-4895c0763f16 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.532496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.532496Z digest=sha256:6c114509366389e39f10a1a4897b787203461d6c47d99b6ee7fa87980574d9f6

Observation 29cab017-bd44-464d-b574-9ca054a70db0 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.725844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.725844Z digest=sha256:f25392a3d2b300bb061c9e474e1228b8fde8dd4383cc0c71aae6f34551b92ab7

Observation bd6dce6e-0cd0-42ae-b33a-bf09de606e61 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Behavior Regularized Offline Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.828145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.828145Z digest=sha256:494de17ea630d9650d28ac7af92b4291f337d2b0e49ce1671ac0a24b0e9573a3

Observation 1db2435d-d8c6-4259-8446-4e8bf1354ca8 · outbound

This paper cites Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.019748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.019748Z digest=sha256:31d6d2435a21f05633a35abd8520dcea6a0d6819b4fcad437df0401aeda26bc1

Observation 5e013ce3-a10a-471b-aa15-3e97c64c1a33 · outbound

This paper cites Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:57:47.046425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:57:46.095320Z digest=sha256:0c43ef0d1b010020b7a0752b7823bf05837b36e7e02c4bb09c84d7352616bab3

Observation e68f0f04-9717-4c55-bae4-336e6ffd1a20 · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.159807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.159807Z digest=sha256:2205165af15aad8be6722b88ff1058f4cf4e5ff9aa41364f978e701cdb5a4508

Observation 526a60c2-719a-4d3d-a26a-1709df4352be · outbound

This paper cites Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.275862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.275862Z digest=sha256:0f9a982b5c1e544ad932ca64a6fe25ea61150bdbda938d6617652a2bc906e2bf

Observation e9dc4928-078a-4b0a-ae05-ac0a84fe3080 · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.353786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.353786Z digest=sha256:e884c48a1cf099d05c99475ee9f32cc787e58fcf6cb62a75b9cd588ab3f32b4c

Observation 1d645912-5e95-4d0f-b754-c3ac3eb5a55b · outbound

This paper cites A closely related problem setting is explored in Jump-start reinforcement learning (JSRL) Uchendu et al.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only A closely related problem setting is explored in Jump-start reinforcement learning (JSRL) Uchendu et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:57:48.185476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:57:46.431032Z digest=sha256:2d51703820da9c106287ec49a8f4f54c532aefdb060e8065cf64eed95cb2b61b

Observation f813b151-7d58-472b-9a56-6ef2d09bd9b0 · outbound

This paper cites to retaining offline data, the approach of not retaining offline data generally results in improved average normalized scores, accompanied by increasing variance.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only to retaining offline data, the approach of not retaining offline data generally results in improved average normalized scores, accompanied by increasing variance

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:57:47.979056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:57:46.516116Z digest=sha256:a8d59475cfb5b63057c43138f7583bc2481373f92a8962246dddc13b8f5c7b6f

Observation ab389c53-f24d-4b5c-915f-931b0bcaea1b · outbound

This paper cites In our experiments, the pre-trained policy serves as the policy prior.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only In our experiments, the pre-trained policy serves as the policy prior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:57:47.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:57:46.595010Z digest=sha256:a4a9b718ce9887975aa78ef1b41d1ca3f6981717bc53b8604e65bcba50ca609b

Observation f62dd1c5-efe2-46b8-bf7e-512d6e575192 · outbound

This paper cites doi: 10.1016/j.robot.2008.10.024.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only doi: 10.1016/j.robot.2008.10.024

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.662518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.662518Z digest=sha256:2e3c1e509a1c4fe6a875e274a8db9cfe2c304cf2ed3b053b4ce04acb3aa9c409

Observation fcd94878-eafd-4482-8672-efd02332bc5d · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.643556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.643556Z digest=sha256:bee461dc726f811a24063cec6b26fb0b53fa9c28f69624644f2ba39fa04b7f0a

Observation 344e76bf-c1c8-4503-b413-d65496c998d9 · outbound

This paper cites Elastic Decision Transformer.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Elastic Decision Transformer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.913949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.913949Z digest=sha256:44a350c28876d9229ac34e2059b9dfec8430a2bd9d59c8db8d23a90208bc8459

Observation 963ee686-3c8d-4410-a165-50f527e98b0b · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.169227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.169227Z digest=sha256:dbef774d4c034144673b5363cdeebcff21fdc19bbce1c0dc9f98d087a98f0d15

Observation 89eecea1-fdcd-4137-960e-c7136a72294c · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Off-policy deep reinforcement learning without exploration

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:57:48.345158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:57:45.091636Z digest=sha256:9505dc5312c84d36de62ad13ab8e0f19be2a9ff1ea4fb82aeae9324fc3ad2587

Observation 20cfc6a2-fa91-4288-86fb-cbf6c81b5c33 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.414360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.414360Z digest=sha256:11f70f43b7a99bc8dd60c76609c3b59e5b13d6e010a8f88f5e8edc3ba08200e3

Observation ef582d84-4c4b-44e8-9edd-05a8791a980d · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.897231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.897231Z digest=sha256:d1969f44a6a66b2714a78b19c444c563ede989be27064f2bd07c320d17e59531

Observation 28005ef9-f6da-4e9e-82e9-5284e244accf · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.786459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.786459Z digest=sha256:5f2dedd676e198e074df7cf8418ad2b975cf9c2f6aef0aceaae4fd6790f70f46

Pith citing papers

Observation 5dc93c48-cbd9-4eda-a6c5-b228f351456b · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.065518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:4f5552651d8b7b936368b06924d6e256d753a88db88131aa99549a4083a242a7