Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:46.595010Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.16856.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:46.595010Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T13:11:16.568415Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T13:13:18.063828Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e7312754-442e-45ba-a88b-ee4ac8fec3f6 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Is Conditional Generative Modeling all you need for Decision-Making?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddec04fc-0c8f-4096-b3bf-0f0d762108ec · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdffa202-d7d7-4237-889c-d54e5cf02b8d · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Offline Reinforcement Learning with Implicit Q-Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1227962e-e897-4bac-8abd-4895c0763f16 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cab017-bd44-464d-b574-9ca054a70db0 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6dce6e-0cd0-42ae-b33a-bf09de606e61 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Behavior Regularized Offline Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db2435d-d8c6-4259-8446-4e8bf1354ca8 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e013ce3-a10a-471b-aa15-3e97c64c1a33 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e68f0f04-9717-4c55-bae4-336e6ffd1a20 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526a60c2-719a-4d3d-a26a-1709df4352be · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9dc4928-078a-4b0a-ae05-ac0a84fe3080 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Reinformer: Max-Return Sequence Modeling for Offline RL
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d645912-5e95-4d0f-b754-c3ac3eb5a55b · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only A closely related problem setting is explored in Jump-start reinforcement learning (JSRL) Uchendu et al
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f813b151-7d58-472b-9a56-6ef2d09bd9b0 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only to retaining offline data, the approach of not retaining offline data generally results in improved average normalized scores, accompanied by increasing variance
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab389c53-f24d-4b5c-915f-931b0bcaea1b · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only In our experiments, the pre-trained policy serves as the policy prior
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f62dd1c5-efe2-46b8-bf7e-512d6e575192 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only doi: 10.1016/j.robot.2008.10.024
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd94878-eafd-4482-8672-efd02332bc5d · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344e76bf-c1c8-4503-b413-d65496c998d9 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Elastic Decision Transformer
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963ee686-3c8d-4410-a165-50f527e98b0b · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only OpenVLA: An Open-Source Vision-Language-Action Model
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89eecea1-fdcd-4137-960e-c7136a72294c · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Off-policy deep reinforcement learning without exploration
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20cfc6a2-fa91-4288-86fb-cbf6c81b5c33 · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef582d84-4c4b-44e8-9edd-05a8791a980d · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28005ef9-f6da-4e9e-82e9-5284e244accf · outbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc93c48-cbd9-4eda-a6c5-b228f351456b · inbound
COOPO: Cyclic Offline-Online Policy Optimization Algorithm Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.