Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2412.07762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07762 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:53:06.334933Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:45.826205Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba3f8c42-beed-4bc0-8510-eae4208a51e4 · inbound

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training cites this paper.

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:06.334933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:53:06.334933Z digest=sha256:3c6ab55f1f42e93827e4563737b3441074e0e9fa94b02341e4ffda899318a580

Observation 010eb31b-5ee2-44fc-adbd-f9db4dea1fe8 · inbound

Reinforcement Learning with Action Chunking cites this paper.

Reinforcement Learning with Action Chunking Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:22:06.470744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T05:18:20.960945Z digest=sha256:66c9328801ad352aea63233e2543e9eceacb1dc1c385a40cbf8e8a20918f501d

Observation 960fc39a-6c78-45a9-8cd2-88d8b797036e · inbound

The Three Regimes of Offline-to-Online Reinforcement Learning cites this paper.

The Three Regimes of Offline-to-Online Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.173421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.173421Z digest=sha256:ab2daee4720a65dc4d944f533d47aa4689e28ca8dbfe84453e5d4dde2f109776

Observation 38c6f8fb-2a80-4507-9e75-29f66546dac2 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:35.297180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:35.297180Z digest=sha256:02e86164c6e78b35538ece9b41ee6eacd76e2ba0ed1487bdf71fc0c5009a99ed

Observation 9e952ef1-391a-48c6-972b-c985f4848809 · inbound

SpikeATac: A Multimodal Tactile Finger with Taxelized Dynamic Sensing for Dexterous Manipulation cites this paper.

SpikeATac: A Multimodal Tactile Finger with Taxelized Dynamic Sensing for Dexterous Manipulation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:51.446369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:51.446369Z digest=sha256:8e04543f099fe841823dd7e6635c8c98fb287064c4fa4e8e573262dd293e7a1b

Observation 80f8f9bf-10fd-4a81-ba9f-f5588dec47e3 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:49:58.825253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T11:48:13.368844Z digest=sha256:db99c650a2b8cba6c3588b470ebe28291c83f7646828a6821911cbec54ecb459

Observation 3f1340c9-6380-40cb-8b8e-a34b9a9ae262 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:55:04.124155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:54:57.866685Z digest=sha256:8f4e71ad86475e4913bb6a08917d779c119011c43fe5e498666c1aac48efcf3b

Observation 9903d06e-5fa8-4ab2-8ab6-764ca48aae88 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.444164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:72d95b54d8b2ca155657d0d23f962a98c9a5009b67c78cc1c87d7e38df1d5ab3

Observation f3a14684-4401-4cdf-a74f-3785b27866ff · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:55.143469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:a7b32609914888be4faa50e2d46b2a469284add73b760a092df0a1c43f19c096

Observation b39afcf6-a0d7-482b-9729-1e050cb9cd33 · inbound

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning cites this paper.

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.325903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:18:05.636131Z digest=sha256:fc2cd0ee1de43c244600e3ebf63bc1b7087579bf3f84ed668c52dfe0c0f2b159

Observation 2b6c922f-9247-4bcc-8025-1453d1d53ec8 · inbound

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation cites this paper.

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:24.122109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:55:12.531822Z digest=sha256:204c8f96ee11f2cce9ffb52f86fff82c846d012842ca93beb533c50668ec387b

Observation 3e7e6bcc-1572-4028-95e5-a15a9bf995d8 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.426634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:8362fb8cf5d24c00345575b2bcfe60a54d846232d793480874b870abe3dd4ab1

Observation 5b4de9fb-0cae-4a8d-a72f-9485efdf5270 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.633737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:8a1e51e9f16d7a37367959b16d81101c6deb5601bcd1003f7c24de0a0e37b3b6

Observation fdd7bff5-439e-4469-9192-caf94b27fbe2 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.827910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:3981ab00142999e814c691ac6ff9f7f8a324bd3aa09efe8ceaf91c62afde2f36

Observation 9c94a5ab-9180-4f85-b548-bbb8d6e8912d · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.054667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:218d37555a331895d2c0a7b9e987ff65f023697424d7228bc1caf7a42a3a3505

Observation cf6d3ac2-d0b0-43a4-b0d5-d4ae7b410429 · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:7e3577e59b076ed50b51442ff740ee1fb67a12075af28ec78658a983bee51996

Observation ebdd9374-8d1a-4aef-bda1-94e5980ee6db · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:24.369491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:24.369491Z digest=sha256:4595c671e402abceb70f982bbb59fb0aa31ef671f545d008b4202a6e852343b3

Observation dce039d1-cbdd-4a12-8533-f5f4580e6e34 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.958427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.958427Z digest=sha256:e5428117e9ccb7d4ff5b8a31a9091823a98bbaf61b3e630e3b2aa027f302226d