Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning via Implicit Imitation Guidance

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 8 inbound Pith citation observations for arXiv:2506.07505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07505 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:46.402936Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:43.858412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T05:12:05.250025Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:98514d551e241c53c2bbbc261adb826aed9aefbc7baadce3131adb6862a39438

Observation c0c7ca42-7fe9-4e31-b0ba-49f239dff920 · outbound

This paper cites epochs per update.

Reinforcement Learning via Implicit Imitation Guidance epochs per update

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.651015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.399321Z digest=sha256:3ee0f521d482cccdc37a342384d841853e663d08de4aaf7a8cb36b63c815cc75

Observation fe3a2452-6743-4417-a163-f222b9a72661 · outbound

This paper cites worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled.

Reinforcement Learning via Implicit Imitation Guidance worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.640381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.402936Z digest=sha256:5b5cf816b5858c3f2cdccc5849d3c73ea74b3f004afdcafd835b6f65a8ddac2a

Observation 532c19c2-6b0f-4037-997e-8780c6c955fc · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Imitation Bootstrapped Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.332349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.332349Z digest=sha256:d2385401fcf0ebecafc02a62cd841e81dda5c8d06e37a997d16ea7810bfc27fb

Observation 06c7343a-0aec-4818-8248-c1f59cbebcfd · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning via Implicit Imitation Guidance Offline Reinforcement Learning with Implicit Q-Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.336559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.336559Z digest=sha256:d2312332a62231f002caac1db6329d7fe2741d7446ed2ef2083c498d32c5357e

Observation b2329927-ac49-4b9c-b574-274bd0fd61ed · outbound

This paper cites Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias.

Reinforcement Learning via Implicit Imitation Guidance Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.344696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.344696Z digest=sha256:21d2e7a3c112ed150a5bfbe68aed9383fccd04aa705764f9985cfd5544c8a2e4

Observation a13b6330-ea3a-4e6e-8793-87762ccd29da · outbound

This paper cites Over- coming exploration in reinforcement learning with demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Over- coming exploration in reinforcement learning with demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.349077Z digest=sha256:d1b339b40b5f54e2df591c71365c2ca86c72a7c6c01aa209b454c9e6a1bca770

Observation 2d76a449-f55e-45b6-9a25-4fe50f74881b · outbound

This paper cites Computational Theories of Curiosity-Driven Learning.

Reinforcement Learning via Implicit Imitation Guidance Computational Theories of Curiosity-Driven Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.541101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.356660Z digest=sha256:cce9b2d25cb93384df171702cb9cb1043d7a509490f6d4a745e051635866c4a9

Observation aeaa12dd-e2cd-41c4-8703-41c978fb7710 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.368469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.368469Z digest=sha256:99defd13f78e8835fc194f26295573fb53280e42611d3a5a7406ee74f0c9a124

Observation 8a60e0dd-1666-43bd-8c84-0e5cc8d313af · outbound

This paper cites Schmidhuber.

Reinforcement Learning via Implicit Imitation Guidance Schmidhuber

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.661903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.376056Z digest=sha256:0023f9d34d59e037287b9007052781884f50fa201cdfcdc0513084953f479b01

Observation 13036d68-8079-4468-a573-ac508a627241 · outbound

This paper cites Parrot: Data-Driven Behavioral Priors for Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Parrot: Data-Driven Behavioral Priors for Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.379573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.379573Z digest=sha256:1c3fabb536885c80d63c6d6a2199fa295bc02bd4673dc48e5f846ce94e8148ab

Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.383535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.383535Z digest=sha256:998649d9defe1145ed866aa1cdaad690f9d66b69da559a3cb18b784bb3468eab

Observation 4474229d-0083-4b59-87ee-9d01368b8b76 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Reinforcement Learning via Implicit Imitation Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.387296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.387296Z digest=sha256:43ce6b73abf327df25c032f4bfebe70fa40aff89e06dcd250d583b1deca1e435

Observation 0066d2d0-370f-462f-b277-9688d2779557 · outbound

This paper cites Learning latent state representation for speeding up exploration.

Reinforcement Learning via Implicit Imitation Guidance Learning latent state representation for speeding up exploration

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.450428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.391312Z digest=sha256:36d066ee13ba5c894fd78985476b2ae905af32df7d04ac5223d9d32f1f82f43b

Observation 5da8b341-7bea-4905-9769-d2ff68f8c0dd · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.395221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.395221Z digest=sha256:cfffd4b8e99eafc1434236c61d263b39ed9992d1321b7863111fe368d24a94d0

Observation 14ab7d25-fb58-442a-b0b2-d46ed3aa6d56 · outbound

This paper cites Exploration by Random Network Distillation.

Reinforcement Learning via Implicit Imitation Guidance Exploration by Random Network Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.319807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.319807Z digest=sha256:1c9eaccec30ed5e3468a54cdaf7f57b5703c320f0a5fb49b47c83a0f986e9797

Observation fd8c01a0-028f-4001-ae53-4883703f0cc9 · outbound

This paper cites Learning by Playing - Solving Sparse Reward Tasks from Scratch.

Reinforcement Learning via Implicit Imitation Guidance Learning by Playing - Solving Sparse Reward Tasks from Scratch

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.496733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:37:46.372311Z digest=sha256:07f9a070f338f31863551edc5312522d3834823e595276c42b7a85b7b747c4ac

Observation 4197caaf-9d9d-41ea-b2dc-43b242038ebc · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Reinforcement Learning via Implicit Imitation Guidance AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.352606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.352606Z digest=sha256:41285d2dc7d492fad47e91b187615809abaaa63f6e61a890d7399dac75656c82

Observation 084fe3ea-9951-45c1-9f97-44b59aa7afea · outbound

This paper cites Self-Supervised Exploration via Disagreement.

Reinforcement Learning via Implicit Imitation Guidance Self-Supervised Exploration via Disagreement

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.364579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.364579Z digest=sha256:40340df2b066c2230ee9d9743260f50a3e28a992385d52bbab7ccfedb0d5c5f0

Observation 18e26509-53b6-4ac0-a579-ec7add5f2be0 · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Go-Explore: a New Approach for Hard-Exploration Problems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.323819Z digest=sha256:7ab6bbf893832bef79d4f8db1ea53309f7aba0998145c2243475d507f265a9d2

Observation 59835236-da50-444b-87cb-09f63dc28f21 · outbound

This paper cites Efficient Exploration via State Marginal Matching.

Reinforcement Learning via Implicit Imitation Guidance Efficient Exploration via State Marginal Matching

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.340669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.340669Z digest=sha256:6ce8ed3bd4101dbfc0e68e36cadc3a0771ad04f26991d432bd22d89399935e9f

Observation 88953db2-4e1a-4aa0-8f4b-bd80c719fd01 · outbound

This paper cites Making Efficient Use of Demonstrations to Solve Hard Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Making Efficient Use of Demonstrations to Solve Hard Exploration Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.360602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.360602Z digest=sha256:4f26d67518cd87632ab7277f66595f45aab3471794121c0b5a6bd531f289fb7e

Observation 88e13b0c-c32f-4da1-9331-d149b31a9721 · outbound

This paper cites MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.328115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.328115Z digest=sha256:6fba81e2e8296df506ce3db3875e10ccdb708650aeff2851b77813dbc57636df

Pith citing papers

Observation f7f7acb0-042e-4fab-b930-ab01711ff06d · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Reinforcement Learning via Implicit Imitation Guidance

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:05.252079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:13d15c5b6b05ab0a87b32ef1275acaeb0ae3d246869244393b3dab30fbcb8009

Observation 912f030f-4650-4d47-b137-548db7d98738 · inbound

Value Flows cites this paper.

Value Flows Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:27.929388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:27.929388Z digest=sha256:f4d77f55298b72c13a31634ea821ab0e0d75f5f11726367853f1b235c0d7395f

Observation 73bb73d3-8468-4d6c-b222-9bd2e736b647 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Reinforcement Learning via Implicit Imitation Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:59:59.505114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:a0ceffc8d40e29c773ddeefa6e225ef2a72341eae6639e0d9b2536f08eb80d65

Observation 740ec016-c10b-4d8c-b869-42282c0b9244 · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.992073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:48:36.372298Z digest=sha256:c4e13298366cd96bb68675ce9f72babe7cf7857875a55a2bf43ec2680cc4544a

Observation 510bd843-417c-4782-93d9-602cb0adf3fb · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:08:53.253285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:08:53.253285Z digest=sha256:716bb15df693ad3ed62287a29899ccf850bc807307ed73a2472e48260c7dacc9

Observation 2eee268b-bfe8-4404-b5b0-805b788a900a · inbound

FASTER: Value-Guided Sampling for Fast RL cites this paper.

FASTER: Value-Guided Sampling for Fast RL Reinforcement Learning via Implicit Imitation Guidance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.234174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:47:36.475845Z digest=sha256:ae295d281dfdd23db39869ea189c4f6ef7a083609ce80bfead290b810ce60978

Observation 6dc67857-44d7-4bbf-85f2-b9e352933bc6 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.032376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.032376Z digest=sha256:f265ac4d8bccc1ac00a6d4a28649b20530d33d1ec9d712295a772e91a5ac4e2d

Observation 49b28c9e-8205-4ea0-b167-2c60e716a353 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.858412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.858412Z digest=sha256:a3a29bf04ff2d4ce93f9c27b0693c54c20d695fedc2a5a410c45eef530b8db32