Pith. sign in

Paper Citation Record · LEDGER

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

As of 22 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2607.27203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27203 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:44.056272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:2c2ac04f30554d90fa1b0cd3f37355fba7c0d64552c4d69bf0d0dff3176dd5ca

Observation 435b5e08-e59e-4b49-8514-ea7445d15ea1 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.855135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.855135Z digest=sha256:3cbf50cc71447b128017355dc060394dd34aec972f5e0719f8960475584ef9b9

Observation 49b28c9e-8205-4ea0-b167-2c60e716a353 · outbound

This paper cites Reinforcement Learning via Implicit Imitation Guidance.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.858412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.858412Z digest=sha256:f788d0268f29fe08d5ddea1da1eb07e18b672014b4362fccf469c6db34b7309f

Observation 3de1e9ba-ad4e-4dc2-8962-641f2c84653a · outbound

This paper cites EXPO: Stable Reinforcement Learning with Expressive Policies.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? EXPO: Stable Reinforcement Learning with Expressive Policies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.862543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.862543Z digest=sha256:bd40627306a2b58294679edc06bb59a71f25e2dd434a0153908f4e0d9f2d7284

Observation b753a510-00a3-4072-b11f-db5da5f4506f · outbound

This paper cites EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:27:44.406212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.865881Z digest=sha256:01dd2fb65c5cc3667c098ba47ef33b633e3975650e1d09c4ded448632859c623

Observation 4f1dd5b7-39d8-4d00-bd48-d23ff9b982ad · outbound

This paper cites Tql: Scaling q-functions with transformers by preventing attention collapse, 2026 b.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Tql: Scaling q-functions with transformers by preventing attention collapse, 2026 b

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.869331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.869331Z digest=sha256:e25cd41caad0b3cc03676539a99fb803d9e757cb7e9055620389f2b206c62a3e

Observation dfb8630c-ab78-40c1-9a6e-8839ca96ed43 · outbound

This paper cites Value Flows.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Value Flows

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.872344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.872344Z digest=sha256:e3f369aae5e1fcf60d6a6ea2a33dc5d87d7535d07acffe8fdf449da7a0be91f5

Observation 1ba61b2f-ad5d-4c3f-90f1-50f313f9c14c · outbound

This paper cites A Minimalist Approach to Offline Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? A Minimalist Approach to Offline Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.875428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.875428Z digest=sha256:950d199012b7acceaeed55f7473b88d470a60b32a9381c7ed7549fa247da46fc

Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · outbound

This paper cites Off-Policy Deep Reinforcement Learning without Exploration.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.878960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.878960Z digest=sha256:e11d0bfe9735ee1592e006c41fa93fd9fe87ab41efc9574c9b1123a10bda427b

Observation e76a3389-5aca-4375-83dd-25a19774da10 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.882699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.882699Z digest=sha256:613505d641635beb2818af082f4c2bb35b64d077a77f29b139d423899be35e45

Observation 30ce3aad-ceb7-4613-88d2-e5bc3e2ef94e · outbound

This paper cites an unresolved cited work.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:27:44.708940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.889168Z digest=sha256:de3bf4010fdf99c722331bbc092f8d648a314673d9fadadaef9d37a0650b22d2

Observation 4ef21e19-4a49-49a0-bd4e-28aec6a260ac · outbound

This paper cites Residual Reinforcement Learning for Robot Control.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Residual Reinforcement Learning for Robot Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.895668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.895668Z digest=sha256:3dcc591983c9606520568f92703ddc4de556e5a65d6f29c9379933dc6dede9d8

Observation 8732657e-48fb-40c0-8468-94f37b81e5f2 · outbound

This paper cites RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.899296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.899296Z digest=sha256:803ad38a87968c687fe1eef0beacb80f32b9badf58ca9a78f7ba4d25f3da5282

Observation 3b1a3d1c-7b99-4f72-a8d7-3ff78b3bd4d5 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.903146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.903146Z digest=sha256:a11bc2e6419b14f35ccb1c936f93e0c5308b37a2a569f52132b49637603f3aef

Observation 58830bf1-7272-4d8e-91da-9c5b9b5759ea · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.909990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.909990Z digest=sha256:bed8ce88af14b45c48bd82a9285b4fdf12a074e7446cbf542815079f5c0b884d

Observation bc29c0e7-e20b-4a77-80af-d7e93fb93bb6 · outbound

This paper cites Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.913078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.913078Z digest=sha256:0a134704897ad563d632ba3f90e24433a4bd1a4f92edc847847e4c1ba5270452

Observation 9ed5d8e6-24f2-4902-a4e4-8c77af3b4e91 · outbound

This paper cites Training language models to follow instructions with human feedback.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.916046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.916046Z digest=sha256:fa8acaf0474ce1c51da1f994bc010049af869e2459ea8c6cffe4b58d7a4daf70

Observation 59243467-6b8b-404f-be6d-410361a58dae · outbound

This paper cites OGBench: Benchmarking Offline Goal-Conditioned RL.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? OGBench: Benchmarking Offline Goal-Conditioned RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.918926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.918926Z digest=sha256:c93e91d94204b5e1399ae61f741573b1e44b69b558101e240433719aa4b0a70e

Observation 56944067-1641-4a97-b8df-a49a69282f82 · outbound

This paper cites Flow Q-Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Flow Q-Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.921979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.921979Z digest=sha256:bd7d99829391d2ac02f29e5e74200cb899243f9db55d20aee39759f44b01b4e5

Observation 5f3bcbff-1be5-4146-9a4e-5b34e7f19672 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.925054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.925054Z digest=sha256:89b467ae148a0f2e288a78b511f8ac703be62afb0ce340d932ea397cc6e1ea0a

Observation 7d59be1e-05dd-4196-8f89-cffe8033cd6b · outbound

This paper cites Diffusion Policy Policy Optimization.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Diffusion Policy Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.928017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.928017Z digest=sha256:c9f32e9bf0d21cdaf11fe70ca587cd466e2f1b375da5772da458b8493ec99f66

Observation c59d1e37-14ea-4bb7-b8ce-d3f0f8f81e93 · outbound

This paper cites Learning from demonstration.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning from demonstration

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.700434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.931195Z digest=sha256:3ee1fecf727213ee8476de85982c46f158763871c6685c3a7a979e027ec81e02

Observation 32783b5f-8530-4ddd-ac87-dceba0a00845 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.933855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.933855Z digest=sha256:34a9d258266856af7327072bd4f44079f62ef320412350a07cdc98c0a197df1b

Observation 63bc5cb3-043d-4cd5-89ed-347195cd729c · outbound

This paper cites Residual Policy Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Residual Policy Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.936921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.936921Z digest=sha256:e0e148fc741dfe4b18cfb4b436e9d307895eb0d318bf2a1597fc005ece5cbb05

Observation cb0d603a-dfea-4cfe-a598-f91434d33dae · outbound

This paper cites Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.940482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.940482Z digest=sha256:9dbf86829ea6938f3b465ebbf03849b30c9630e9abf3ede1f942110d38e9e55c

Observation 3bb83959-8101-4961-9f60-cad24838613f · outbound

This paper cites Jump-start reinforcement learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Jump-start reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.691778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.943635Z digest=sha256:b44199591697b709808afb42675dfd780791f744a32036eb73ac25f40d5fb25a

Observation 5cb8093f-1c55-45d5-b11b-e15ab67eeaa8 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.946235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.946235Z digest=sha256:9e39556468246a5fc11bcd5bffbf019ad52eb6d73a51d7c2821046b5a80da6a8

Observation b9a10551-8f68-4ff2-845b-e4b2c0c6343a · outbound

This paper cites Posterior behavioral cloning: Pretraining bc policies for efficient rl finetuning, 2025.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Posterior behavioral cloning: Pretraining bc policies for efficient rl finetuning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.949288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.949288Z digest=sha256:638e6c840256c58f213af9ac072bbf823a9c98191588ba218dd8c87ae9c04141

Observation 927ceae2-6df2-4eb7-b7a2-05f2ce14d0a3 · outbound

This paper cites Hybrid policy optimization from imperfect demonstrations.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Hybrid policy optimization from imperfect demonstrations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.682421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.952021Z digest=sha256:8e310193f1b7395a44fad844e04ddde81f8230fa5b631879328af2850675b8b3

Observation 94deceb8-0c11-4433-b87a-394c4ee5677b · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.955043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.955043Z digest=sha256:aab68aaaf01bbdbadf796f145e6cd9528676a4a5ef848ffa936a8a4d107c4e6f

Observation dce039d1-cbdd-4a12-8533-f5f4580e6e34 · outbound

This paper cites Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.958427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.958427Z digest=sha256:5065a444bb04966dc6ddb56573af85ff0adde261953f8d0a2bec44dd9d576503

Observation 1066cbd0-4f29-44c3-9c82-b26fb9b31e4c · outbound

This paper cites 2024 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2024 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.673796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.961450Z digest=sha256:32c77e64fee10a8cd2fa8fb8895a273ae6781027945036cdd411c11200bd4358

Observation 7835ab6d-fdf0-4e5a-af30-89b276e89431 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.665295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.964400Z digest=sha256:9e5a43f0f5e24810fa1956a3e76118774745ba94d4e15babdc3762a3898de934

Observation 78f74e7c-1372-412e-805e-151b93895833 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.656865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.967979Z digest=sha256:cb9703bef2658bc5a437bf8cbdf9549e830988769d814797d467a6a76b436455

Observation ff6ca7d4-31ee-4b3f-86c9-dae97b4d45b1 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.970826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.970826Z digest=sha256:6ce9539fc7b0462ad1136c8a2bed02966399e87366f7634ec60ba5d5723435dd

Observation 39bfc212-b23c-4369-ad9b-b4b6baa6585e · outbound

This paper cites Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.973679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.973679Z digest=sha256:e5cd7b4856ac5bc2aae2aba7d0dd5e9361d3a578071832a63cf170bc000790d2

Observation 8694ecfc-f619-4c79-bc33-0cc52e54f1d3 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.641172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.976727Z digest=sha256:d41000149d9a9991b0db1f933f5d40629bfe89c8f9a3b363c684d6583473358f

Observation 0d33480c-c323-4ef5-881f-ccfe952540f4 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.632942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.979698Z digest=sha256:234f1a50b7fe7258ef24427e36990cabe4dffdb63fd742bd95e1ff02bb9bcdc3

Observation e789b70c-4153-4194-9bf6-a7c476d0876e · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.624224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.983043Z digest=sha256:54630bc4cf6de4e568bd7f4647b5cad36a7d4acbef3f6ccadd1766f706725deb

Observation d1980857-ee60-4eba-9aa3-c30134053f99 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.615817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.985785Z digest=sha256:1c3e07fa810584466306500b9952f9c9b070e848e69f2e10b270e66936be9762

Observation dcf5a4ef-cd04-41e1-9252-af049d922457 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.988561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.988561Z digest=sha256:df2f089471e8018bb5c120f0a47d164c7106863e77fb4f9d296d79394b51eff1

Observation d2787269-2f10-4c59-a4c2-3dec0ccc7b6f · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.602667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.991903Z digest=sha256:daee51da2cd70ea452a54a40ec78215e17f1380b3e6d643d29638e40df7f3648

Observation 6e3502c1-8b5f-4acf-8f1a-3b3868634f40 · outbound

This paper cites 2022 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2022 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.994697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.994697Z digest=sha256:822a8336c5e6ee57c0aa7367bdd82247622e0d5c8289331e935ba091bad9e7a5

Observation 261c0147-f982-4990-b942-df7184291f2e · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.997602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.997602Z digest=sha256:af23348ee71b026f8b97914477c4f9c28575b9252fb857d42aaba85b86c21e50

Observation 96736635-cc01-4463-a940-97b4f349a94e · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.589458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.000430Z digest=sha256:a34f775633404c027424b4d2b8687884f18ef8d061735c188af52a981c16722f

Observation e588344d-82c4-4db6-9258-71db439f99c9 · outbound

This paper cites 2019 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2019 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.004064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.004064Z digest=sha256:a68d014f6f7cfe667418724d8b4779b7b9889fda12763ca0ce3ef377f960ad0f

Observation 3340c488-d9d8-4969-8f98-e65a4360125b · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.575786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.006852Z digest=sha256:2269be0eef6e951aecd9fe13eb879810221baf761b4fd53f56098a2399375484

Observation 6e818da6-7c3f-4ff9-85f4-d8474843bbfa · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.566579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.009721Z digest=sha256:5c9d0ff8eb3a141ac160999a1434b3ef309e237cbfe74cf327d796ee2d43aac1

Observation ff2fc180-aa2b-46f3-ae96-a689cd150192 · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.558137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.012520Z digest=sha256:e0f548d71037ceb030b39fa7b77155968e611e73f9f9fc4af2380e350260387c

Observation abc7736b-dbb5-47b9-a59f-a4c8830b68ad · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.549631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.015214Z digest=sha256:2f8033dd192ebf27b745effae1884d7845c2e01b4614dd1ca3fe3cf05aa85d4a

Observation 8e547edc-336b-4161-8db1-bd7993eca08b · outbound

This paper cites 2019 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2019 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.541037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.018125Z digest=sha256:fa5d1867cb055a7424dc6b1f08c4780f49769ea4561888d45537557a8bccf50f

Observation d034223c-a59b-4923-ac80-a71d0870a309 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.532567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.021235Z digest=sha256:71e90147f0a1935b210e27b6cb00300acf42287e9443cd23e4b9de26625690f6

Observation c8d85fc6-f76f-49cd-a26c-0f1333988e7a · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.023977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.023977Z digest=sha256:5a13691b9cfc99be553a654c707e5d2911c15a8510e799b3768c894006e857ae

Observation 47dd6549-d3ae-4431-af91-e43b699b93cc · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Imitation Bootstrapped Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.027011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.027011Z digest=sha256:4d8eec12f445b842823efe1e4889e320134a41398235edec949bfa99e1a5726b

Observation 3f7ed219-f687-47a8-997d-2073b0fb3a45 · outbound

This paper cites 2017 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2017 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.029744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.029744Z digest=sha256:3c83d03baef43558fef6b49c60b0ba9b6592bc8e62cc3ec4708c53c383fe53d3

Observation b2d80ac2-af79-401f-8317-cec902a0e667 · outbound

This paper cites Learning from Demonstration , url =.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning from Demonstration , url =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.513790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.032496Z digest=sha256:cff2dfdf99b7678441fd81c9e9c2310374d15a3de5a99fadfd26f85b8a365bc3

Observation a4a62e94-6014-4bc9-8471-5a6f91f827f0 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.505056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.035984Z digest=sha256:730dbafe58f3e107c50963f7c6079e942bbf02ce7621a3da74bf6d3363f46d6f

Observation c5580547-0659-4f5a-bcce-2ef08c83c0a9 · outbound

This paper cites 2024 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.038696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.038696Z digest=sha256:501233dbf30c757e0bed3981d2eb981d98de149ca47def211f606d76347c5029

Observation 388ee01e-ae08-48d0-a270-2a8eb908f0e5 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.490436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.041526Z digest=sha256:2f64bcc1ecf48c07053d8d732ce0a1d78d6c101d14d9b13a5cd7b926774769ea

Observation f4d91b64-1675-4b94-a1a9-8fd81f082c41 · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.480664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.044253Z digest=sha256:8a64d1850ee1af7648c851d522c37876ffa24b48f27c4b3c330d120337f287c7

Observation b68ff781-7502-41e6-912b-c022c4320ba4 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.471999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.047941Z digest=sha256:5cb6c7e9072a6a473519a2594b8eb547cad957a2012624f559463a0fe89a271a

Observation c548dbad-505e-48e7-9df4-4407752df131 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.462445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.050890Z digest=sha256:da4d3c8adc3a8476643ded9b6c5bfc1b723ccaf9263c4a0ba43d30ce49c0b3df

Observation 494a7ca2-e0ed-46dc-8428-a1cd07b0dfa0 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.453888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.053550Z digest=sha256:db92557aa2b9386c2a3f7b9b9b5cb0ff978d3239e31474a16a3389787356d1af

Observation 48b63200-b705-4a67-beb3-e8378cfcfaf2 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.445380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.056272Z digest=sha256:e707dec72ba0119fcc3bd164ed5c584b51ae6dafbe6a946c3b954da791fcf1eb

Pith citing papers

No inbound Pith citation observations are available.