Pith. sign in

Paper Citation Record · LEDGER

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

As of 20 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 31 inbound Pith citation observations for arXiv:2412.06685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06685 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:28:38.548957Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:23:38.206985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.649113Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3682650-6a4a-4d7c-a290-040be540b661 · outbound

This paper cites Abdolmaleki, J.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Abdolmaleki, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.762199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.777772Z digest=sha256:b08b494dce4238fa8011990da40dae928a1823e686efedd1d683387c69b9744b

Observation 57210eb1-0b26-4073-b9cd-de722e8ef3cd · outbound

This paper cites DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.783068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.783068Z digest=sha256:144bf4dcbc58c491e12cf158ade6782574322e5517c3b3ab05db40be1bc882b5

Observation 3756bff2-2cb0-4f3d-bbd5-ab83f380eccc · outbound

This paper cites Efficient online reinforcement learning with offline data.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Efficient online reinforcement learning with offline data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.788366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.788366Z digest=sha256:ff4f158c414e7d177c2cbdcbad42a8afca603ceb128d79c7d3ea303aa8436fd0

Observation c640ceba-ea63-478d-9769-dc38c584847c · outbound

This paper cites Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.793143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.793143Z digest=sha256:80e7fca15d0f5c7d89da3446d913f96d474f016fde863b475fae60671cb3ab10

Observation 85351a09-9f2b-4bdf-9146-2437a62c1183 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone FireAct: Toward Language Agent Fine-tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.798320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.798320Z digest=sha256:43be55520560f8c117ecddaaacec32c6c50ad36dc1ad5d135d976ac22777fd1b

Observation a9c44608-52ce-4d84-8189-988d0b3f0e39 · outbound

This paper cites Diffusionpolicy: Visuomotorpolicylearningviaactiondiffusion.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusionpolicy: Visuomotorpolicylearningviaactiondiffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.638278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.803294Z digest=sha256:ba4b0afb9b5bbf9ae8348c1feb01356e0728db1d77b68e42da5e4b9ae0fb3ba8

Observation 7f289261-15e4-4165-a12b-3283755a8f5b · outbound

This paper cites Bridge data: Boosting generalization of robotic skills with cross-domain datasets.Robotics: Science and Systems, 2022.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridge data: Boosting generalization of robotic skills with cross-domain datasets.Robotics: Science and Systems, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.503358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.808438Z digest=sha256:5104fdd1a29fc7c51b371324ab951bc0c82d8c356b8b14f89398646d7249dc04

Observation cea8ecc9-0983-428c-ba04-b08daa2027e2 · outbound

This paper cites Stop regressing: Training value 16 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone functions via classification for scalable deep rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Stop regressing: Training value 16 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone functions via classification for scalable deep rl

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.332367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.813070Z digest=sha256:ace109315029c3334e79a0ad925b855fdf1839c75c6f8907ae851cf809941bbb

Observation fe6e76d3-abec-4f35-a97c-05ea505deb62 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.818160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.818160Z digest=sha256:886645211d5349c761f7532d8d2d2f458e28f5a29c6f3df73ff3fc00ad5f58d5

Observation 4b8af79e-74b6-4c28-8c47-cd9e726ab9b7 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone A minimalist approach to offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.822367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.822367Z digest=sha256:5242d01a81c593f8c575c20bbb1da40d1f0d93ab5f49f55a16cb5f1dd58a226f

Observation 86c6269a-6684-48ff-aac8-9c76364803f2 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Addressing function approximation error in actor-critic methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.827415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.827415Z digest=sha256:01dd9ea872c2a79a7716cc809192119ac4954305961d078c2aa94a43a2b93ed8

Observation f7603abf-cff8-4374-bcc1-f8de43aa4872 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.832515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.832515Z digest=sha256:6d87b1a175652f760a404169c6f2c0c13a2ada786e75174f01082050eab8aea3

Observation 33f97c0a-d888-424c-84bd-df3892676277 · outbound

This paper cites Emaq: Expected- max q-learning operator for simple yet effective offline and online rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Emaq: Expected- max q-learning operator for simple yet effective offline and online rl

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.285717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.836857Z digest=sha256:b4cb6efc37f83d1af1a188ee78534da21932a3ff935ac58d02ec2f9b35465c40

Observation dfeb17e4-84a0-4cd7-b97a-b32763da0934 · outbound

This paper cites Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.269847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.841410Z digest=sha256:8dd3720da3145a85822c74491d1a080d110523391f3dfeb7584023e13bdc9795

Observation 9e049b36-5eb2-4ef9-8322-49db2d15f830 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.846301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.846301Z digest=sha256:27d150525bf3dcfe881b7c524af637a507824c00559c784b3cb9a76c84ce9c45

Observation fcd3e3b7-9c83-4ba5-b143-31e20c2cc21c · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.851080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.851080Z digest=sha256:68d56abf61753281e020cba042c6cf1dc7b40f4db1e2c213dfdc4c6367ce4823

Observation 1419ef9a-d790-446c-9adb-562c0f9be061 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.855741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.855741Z digest=sha256:53b7f0acf0123bfe63264978311314b57681504598516ebeffeedf3b8d7e5dbc

Observation 900932f6-9fa5-4bda-80de-f8db9fe56cad · outbound

This paper cites Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.896663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.896663Z digest=sha256:1897ff078ee6df531fbd7b35fa12d41e01e99770655e67953e63bb9de680acfa

Observation 8c8718a4-beaf-486d-9b04-86ee4a729b54 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.935059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.935059Z digest=sha256:12ac1106e3a3a90c62713eb6094d6ee2978ab46556c32f05073a6b5a47f32b71

Observation 35a0a684-f0c9-47c0-8fd4-84020e7182dc · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning as one big sequence modeling problem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.213089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:37.975759Z digest=sha256:93839a498492c64d3e0b3dc4ad3f5bf0ac0a31bea7e8659d9035c8af08a9e25b

Observation 7ac1b492-06d7-4e66-9575-0adf6385fb92 · outbound

This paper cites Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.126016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.015147Z digest=sha256:e2aa8c296659a6a5f57b691213e6e00581c3d2980feb005594fcdf1522ebe2cd

Observation 1f6e5670-a8d5-48f2-ac50-1c6e6ab01acf · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.085466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.085466Z digest=sha256:9bd8d36d2ddaaa50bdca8e92e412ed4e62c6ce99b671faba3c5898f7af0c62fc

Observation 393c17c9-5904-4e36-bf7a-a34d0973a03f · outbound

This paper cites Adam: A method for stochastic optimization.International Conference on Learning Representations (ICLR), 2015.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Adam: A method for stochastic optimization.International Conference on Learning Representations (ICLR), 2015

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.110467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.151870Z digest=sha256:58e4c68528f52d11d41339f299aadb93ad20de6df2f623705cf9ab7679b818c9

Observation 6e224e8b-02cc-4e5e-ac33-cc3e93572b68 · outbound

This paper cites Offline reinforcement learning with implicit q- learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning with implicit q- learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.095697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.156808Z digest=sha256:e593381cfed5d5a98481331db048746280725d50313e7b3ad774f56bc055c3a5

Observation 79359524-3e4b-46fc-b2d5-2a8a0875a164 · outbound

This paper cites Kumar, X.B.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Kumar, X.B

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.077671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.162001Z digest=sha256:94688aaf54d1fc3098451e244cd8139c43d12a62328294869a67d7f3af06d06b

Observation 775538a6-0dc3-460d-9935-cbb9cf9f8c34 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33:1179–1191, 2020.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33:1179–1191, 2020

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.166784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.166784Z digest=sha256:23a597019464567a765afdf5a01fe853fdc5a2991b7dfc05eb46525efa9de5f4

Observation c8ea5083-aa4e-4af1-98d5-343aae6bfc45 · outbound

This paper cites Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.172066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.172066Z digest=sha256:649f1f9c58f0ebd04e20445d57e559862e6eb907fadc2b11cbe920be06f9e054

Observation e3a74265-7d7d-4d62-b4b4-a1440a509201 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.176724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.176724Z digest=sha256:ca515c1df38c463b2bc0711797f1865c7ec75ecae2e796ddab233b58c84b2920

Observation b542df0b-9d74-405b-a4fb-5849c2da2e5d · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.025368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.181708Z digest=sha256:7746770fff7c10caf82004dd7cf1b3ac66345f58447a9c1e1db5b8b8b2aa638e

Observation 6be748c2-c87b-41ba-a477-ed2861044d62 · outbound

This paper cites Continuous control with deep reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Continuous control with deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.185750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.185750Z digest=sha256:cc9b0d56814a90553f9c5ed776385c20e00860d9c133f942065cd7c55150240d

Observation ccef55f1-8c98-40b7-8b32-6d2f4de27bbd · outbound

This paper cites Leveraging exploration in off-policy algorithms via normalizing flows.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Leveraging exploration in off-policy algorithms via normalizing flows

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.915108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.190192Z digest=sha256:18d8d1378fa994dfc4f71c651613ebdafcf41b274915f2e9e822132a61522758

Observation f33dd6af-30ba-42af-b754-8a1c475586e0 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.194784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.194784Z digest=sha256:d493dbf1a10727121bb42f681ab2ebefa632dd497537635a712ebdaf7974b5f1

Observation 81a73538-74a0-4250-8201-8018fa5efa2e · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.198810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.198810Z digest=sha256:8f08c626605f331c8d357b4587ef33be705c9a4aa8c56357878f86f46f8a9c75

Observation 8658d600-0fce-4e4f-9475-4523f84572db · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.Conference on Robot Learning (CoRL), 2024.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Steering your generalists: Improving robotic foundation models via value guidance.Conference on Robot Learning (CoRL), 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.827158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.203147Z digest=sha256:ad6816666d5eeb489e3a2126ceb3a90bfad8f2fd7189c8f7e5918e95adf8a279

Observation e0c1692d-7dc3-4467-b535-c87bef76b780 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.811700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.207403Z digest=sha256:722af635ecea7b5b91048fc9a5af348879fefc3120c0e2e823e556a9f16a4bac

Observation d113f657-f8b1-4641-86b1-34531597c672 · outbound

This paper cites Greedy actor-critic: A new conditional cross-entropy method for policy improvement.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Greedy actor-critic: A new conditional cross-entropy method for policy improvement

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.795425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.212797Z digest=sha256:d9b86c8211b78cd938e32253260e4bc1aeb37fcdf20f04f2bbc4016716eaa6af

Observation b2e2af66-2abf-4477-8231-f1dd1d53af56 · outbound

This paper cites Self-imitation learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Self-imitation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.778973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.217614Z digest=sha256:1768716a8cc3a2581a364a55c33c30183b15709908b0b707ab663af4c2cdc214

Observation 081fa1d1-61be-4a29-b247-bedcb98d486d · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.222122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.222122Z digest=sha256:51e4dc2eb49b88370b1c4b24f64e801ba73dbfaa317082ed76b8338168338e52

Observation 5bc20772-9a24-4b5c-b8bd-184f68796180 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.226671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.226671Z digest=sha256:88f114088f95047ec2c4e5a810ee042dcec3d8ad5d546e0f80714fa0dd4662d8

Observation 1d9f211b-d82d-4b83-b548-731c999f232d · outbound

This paper cites Peters and S.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Peters and S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.647040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.230916Z digest=sha256:efa3e5da4483af5e0094ec0617c3ba06fa01a091c2d20d394444e33db1f03618

Observation a24d8794-b64c-4072-ac60-d338d34c7b5f · outbound

This paper cites Relative entropy policy search.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relative entropy policy search

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.566816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.235007Z digest=sha256:2f7f6511d16a41990fa9ae0ca59d5b54aee6df9b136e468e402091077cab9e6b

Observation baeddc20-2273-4533-9e6f-f4778a91a6f2 · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning a diffusion model policy from rewards via q-score matching

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.550340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.239427Z digest=sha256:9a231c4292482ed7eb9bb236db142e46620631204683157f41ed9a27b27bfb75

Observation 56fe718b-1dd6-47f2-9a18-cc015507edaf · outbound

This paper cites Diffusion Policy Policy Optimization.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion Policy Policy Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.244395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.244395Z digest=sha256:7b5a24265fbd234667f5fab1dd99f388ae249836fa5433dace2542ac3fb5c985

Observation dc0a1f4b-0a67-4d01-a0b4-770a9365f531 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.249232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.249232Z digest=sha256:45499b04336da315c9ccc80b50d43dc33513dbcb8c75c04b675ea418c30cc5ab

Observation 79a038d1-d0e9-46fc-adb9-92ee0820caa6 · outbound

This paper cites Grac: Self- guided and self-regularized actor-critic.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Grac: Self- guided and self-regularized actor-critic

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.326396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.257838Z digest=sha256:9eab011eefba85ad569f806473aa3b522fb7b0af3df402da81cdd93af5555dd8

Observation 221685e3-8468-4755-a020-05eddb470901 · outbound

This paper cites Skill-based model-based reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Skill-based model-based reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.302833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.332453Z digest=sha256:fbe4c5f8c2e30d245f9b709d3bffc3db7ec604c1f9e1d8d31b3d7a3465f36e29

Observation d14fd665-77f4-409e-8587-75aaf4caee95 · outbound

This paper cites Q-Learning for Continuous Actions with Cross-Entropy Guided Policies.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-Learning for Continuous Actions with Cross-Entropy Guided Policies

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.396922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.396922Z digest=sha256:45242d9cfa47f90485436b9194fcb6f2940e3f4f4981a67afc94eb8331ca0b76

Observation a22cd976-2b68-4d26-be2f-644f254f58bc · outbound

This paper cites Hybrid RL:UsingbothofflineandonlinedatacanmakeRLefficient.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Hybrid RL:UsingbothofflineandonlinedatacanmakeRLefficient

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.284548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.459011Z digest=sha256:c56cc767fb755c0f2c4528d9d159ad4927c900780577df837dd936282ded0bd2

Observation 24994787-2415-4d81-a67c-61527fce5091 · outbound

This paper cites Second edition, 2018.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Second edition, 2018

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.501183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.501183Z digest=sha256:3bbb9c73c34f41757ff6e853926f44275db8ad3da6d993675c314007d72b4f7c

Observation b74b0c52-2c29-4f23-813d-c70d0a960b5d · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.256265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.505556Z digest=sha256:462ee647a26096f617a4b78d2e73ce8c0bf6d50f6101506180cec8cdad1cba71

Observation ddd37579-adeb-407f-a972-a93d3dffe64c · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridgedata v2: A dataset for robot learning at scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.510594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.510594Z digest=sha256:785369dd6f7c3fda013950ca5f3dd92a88c54ca5e1319d2e02dfc1dcbc3302f4

Observation 0b474557-84da-40f9-a582-bb79fa41abf9 · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.515332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.515332Z digest=sha256:e4da1c0930603975ad43e0dd25356402804c82ac7b6b8513e98114ca7cebbe8e

Observation 61a61b21-4098-445b-a622-ebf21a549ae2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.519972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.519972Z digest=sha256:a123f86278fd6fc60c37e41a3a1e62af9e3397d968518842e6fa9823ffbccc77

Observation 5ed0b98b-a5f9-42a4-a95c-3ce9e28b4a5f · outbound

This paper cites V-former: Offline RL with temporally-extended actions, 2024.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone V-former: Offline RL with temporally-extended actions, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.139198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.524739Z digest=sha256:a2847f055168630aebc097a53960bef3bc5f6f64ebc682aae772befef89cbdaa

Observation 3a02aeaf-8436-4d36-8fe7-741fe74598f3 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.974717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.529793Z digest=sha256:952fc8b152c3bbe656322d2c81f227a99891a2fa2419ce35fe147f464582f2ad

Observation 88d1e079-8e37-4248-a4c5-d907d1663a40 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.534595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.534595Z digest=sha256:59a33990652bba086382381517875ee1bfaf92f8cd4fd04674d2c4fe69ca9485

Observation 98f2ede7-b571-46b7-9152-e41df897f2bd · outbound

This paper cites Mastering visual continuous control: Improved data-augmented reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Mastering visual continuous control: Improved data-augmented reinforcement learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.904003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.539705Z digest=sha256:f18ce009afcb0118ad93586c52ee7dee50fea232bdda7f24152761e0ca43947e

Observation 739b6e11-9382-4951-b6b5-b9243dde2068 · outbound

This paper cites Autonomous improvement of instruction following skills via foundation models.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Autonomous improvement of instruction following skills via foundation models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.887226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.544436Z digest=sha256:41f87f0564b1bce4c914363284f9e66532798c4882530b4c98e7b59f82edc917

Observation d7aaaa09-ce8a-477e-8bb3-68901f58d8f2 · outbound

This paper cites -v0” antmaze datasets from D4RL, but Fu et al.[9] deprecated the “-v0.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone -v0” antmaze datasets from D4RL, but Fu et al.[9] deprecated the “-v0

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.870057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:28:38.548957Z digest=sha256:4a7c91598cc938afc5090e3ba5e8740ff6ffcf9feaf23bd6f0185f748c50c0d7

Pith citing papers

Observation 8c7411f3-fee2-422d-9eaf-f55ce5b57aa2 · inbound

Flow Q-Learning cites this paper.

Flow Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:10.467494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:54:10.467494Z digest=sha256:9425f77d896b7d58800011d007c386ba6a24b379375a541879995cb9d1845782

Observation 130da234-bcdd-4f93-8ff5-e34a16e1b8b0 · inbound

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy cites this paper.

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:20:33.596522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:20:33.596522Z digest=sha256:1e9e53d62407aeb340d7e284078e0618c74e1674ed26f8347d0fac50c93fdcc6

Observation 4c5de78b-2d75-439a-83d5-9af270e71cc1 · inbound

Exploratory Diffusion Model for Unsupervised Reinforcement Learning cites this paper.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.339718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.339718Z digest=sha256:dc1d23f0e79f7d641e19b6d092b03952a4e707cde38fa8ae38dc04d94d851aaf

Observation 0ad9fb00-7442-47ff-b970-fb7420bd510e · inbound

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning cites this paper.

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:38.206985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:23:38.206985Z digest=sha256:6c13369ff61820df6814dcfc00c569c9d296eb95f4f1d5f07906715599525f7f

Observation a5446eea-762d-424c-bc91-e942572d7a78 · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.612205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.612205Z digest=sha256:d7f0235eb338d7b6b898f8cfc6345189257ce84e7d139f2e536befceff7f3650

Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · inbound

Diffusion Guidance Is a Controllable Policy Improvement Operator cites this paper.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.792421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.792421Z digest=sha256:fdaf4d4098f8a24b9ca9eea50ef701869d172b7ef52f104a732e56a6763dfb07

Observation 4335a470-040a-42bf-9ca8-fbf1a0758d2f · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.258768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:0db5c4b0614d74344a433f9fea8e139959ed9cfef206e42607a93b47d5cf8f34

Observation cdd8c86c-0b3b-4f1a-af6f-f9b4914be986 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.611023Z digest=sha256:9738d5889730504e193cdf65b63be011ceb26238c49393ea5cbca8cc34a7749c

Observation 609b4a33-52c6-4b10-aecf-d739fcbb9438 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.359504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ef2bf1c1ebbc69926ad0b12c01e65179514415eb4092e2184c6107f6d5892636

Observation 88261a3b-d278-416d-8adc-03952f33d58e · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:26.050159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:ce145c45fe62e3810ff4746c89de0c99d352bfb3a58af1cb7703a2840946f4d2

Observation 6e7b3928-7f87-4e28-8f30-79961d2b7c67 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:49:58.836520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T11:48:13.368844Z digest=sha256:b2f01dc02d79d2a85279fd3bbdd9c2b262c3fea832dc0ef6a6e27f97d8adc75f

Observation 12e3cd77-d007-4137-8d44-38141bb4cb74 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:55:04.159093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T11:54:57.866685Z digest=sha256:32578474ed6d27e1e1e478c9643e66f268af7d32097585fb3137b99db4a644ee

Observation 00521669-6dec-42f6-945a-8f2473994a8e · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:998a1ed5e3de9544a957900ea9e382faa11dff80bee2e38f4b31c39c61ee9736

Observation e59a8f7b-b208-4f10-bd55-49962d8816da · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.630018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:b57eab08938d065278ec3d6a51d7ff780a401f78e81ef79d9d1c1c01905b023e

Observation dd0e3994-6bfe-48ab-98aa-0a0d3fe5c70d · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.480995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:19b28089514079fe472a8b1884bece434bc3e9b5edd497b3b404796e3dfc261a

Observation c6b3ca7a-5ffb-44a1-9f66-0bf9626af0c8 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.674048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d6e5beb306945e3b909e2e9f1422c7b84e227460c537b9f86ea31eb002a07419

Observation eaafebcc-fb68-4ffd-9c74-5bdd74bfe522 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:53.258341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:d9630168875a5d22656b07a4cfc453ee6d3d40730dc318acb0e1d5b66ea3586d

Observation f5d912b0-9f41-46d9-acd5-ea7348a7f017 · inbound

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding cites this paper.

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:56.223208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:21:23.325806Z digest=sha256:f3f52cde3329b002d9ecaf1eb5a0946cb7023c7c5ad8885265dad6fb9147216a

Observation 06a164a0-6f2d-4115-bdb6-49fcb209186a · inbound

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding cites this paper.

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.522744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T07:41:12.038619Z digest=sha256:c14f7f3377f4eeec9a3766f52042d578b74c11aa632c267d290addab0ce89db3

Observation a8263df0-2b6d-435c-8b68-aa9a8721d04d · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.469693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:366cca988b9e753d16e30d403dd4ec912660e7a1d66d64c091eddaffff07c301

Observation 72bf615c-edb6-42d7-bd65-6b99f9a55500 · inbound

Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning cites this paper.

Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:04:43.127249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T08:02:27.951445Z digest=sha256:641b3418d97b359786246642d1906a069b347a583777a746237f21c42c986586

Observation 7eb4df5d-9408-47fd-b34e-d984144a8237 · inbound

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models cites this paper.

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.959872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:10:08.682307Z digest=sha256:aebe6e9352ec2d6efaf897d4e0924b0fcc95c19d71f17fd23e37f583f0bd981d

Observation 15bae689-3df4-4612-b061-054739eb52de · inbound

MODIP: Efficient Model-Based Optimization for Diffusion Policies cites this paper.

MODIP: Efficient Model-Based Optimization for Diffusion Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.560909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:44:19.708550Z digest=sha256:8e39b0f81a07c756f91808e8709532008870f2261acf177a2cd2c2b449e5df78

Observation 2126534e-497d-41c7-bdb2-2c6513451b7d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.921976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:47afcfbcb1fa263680ecb57ae559eb47f18f38b1b84fc07a5ec1cd7a70f3983d

Observation 5e6891cb-20f9-4ec3-972d-18881248cdac · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.726202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:63cc69d18061a62c0c1a4b715a84d5b7c7d724485d7ace91734d841edb15bebe

Observation bf5b6a4e-37b3-4eb4-9983-2349a56664a5 · inbound

DiPOD: Diffusion Policy Optimization without Drifting Apart cites this paper.

DiPOD: Diffusion Policy Optimization without Drifting Apart Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.948665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:3c7a2e2023add4736dd240d9e8136186f9bc27ea127b5b98ac2e40a73975be63

Observation d861aa65-49e0-4a40-a343-920ed4c51fe0 · inbound

Reversal Q-Learning cites this paper.

Reversal Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T18:38:49.871380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T02:30:24.951689Z digest=sha256:9682bdc9784d9dd17782e502f2f222633d17d4bcbd908143cefc44a1cef06e6f

Observation b93a7f98-09d0-4631-8508-e0dd233a94dc · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.650888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:d07ac4e1df056270ac742215fae44b3349ca4db90755055e200fe2d92894dc94

Observation 9733d33a-8d5b-4d57-9a76-cd74781f6036 · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.657378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:3a7c2cc7259e75b0db3ea6595cedb83284b052bb798e1d2aeff0bbf48b101229

Observation 9f8c20cc-bc8e-4874-bb85-8fbe4589b096 · inbound

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models cites this paper.

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T00:12:42.173815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:12:42.173815Z digest=sha256:ba5cd3e203e0225f769ec6c529c7dd6db0e576464a15bd709e5fc247d31e4dbc

Observation ade8d8dd-59e8-4d12-b661-e9019c6d7f85 · inbound

Adaptation of Generalist Robot Policies with Minimal Data cites this paper.

Adaptation of Generalist Robot Policies with Minimal Data Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:18:48.614940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:18:48.614940Z digest=sha256:0f72835159bd77f3e2cd5429aaa9463a141e34e3fe88c993553456fb8f6b0af8