Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning via Implicit Imitation Guidance

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 8 inbound Pith citation observations for arXiv:2506.07505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07505 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:46.402936Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:43.858412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T05:12:05.250025Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:fd25f0699d5a989b40c914a5b6f540f10426334638dad7d35c991a8bf919037a

Observation c0c7ca42-7fe9-4e31-b0ba-49f239dff920 · outbound

This paper cites epochs per update.

Reinforcement Learning via Implicit Imitation Guidance epochs per update

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.651015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.399321Z digest=sha256:c138cf9fdb1513f8b30c749dd4ae45bfc592dda484b2457cce91cb90f09a1eb0

Observation fe3a2452-6743-4417-a163-f222b9a72661 · outbound

This paper cites worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled.

Reinforcement Learning via Implicit Imitation Guidance worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.640381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.402936Z digest=sha256:3ad9dbfd2c32b16ce7e7b59e5f74c3a823694553e960dfbe66ee108b646bc818

Observation 532c19c2-6b0f-4037-997e-8780c6c955fc · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Imitation Bootstrapped Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.332349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.332349Z digest=sha256:4ad4f1c9909d4e014ffe679dd2f988ffbcca72eb7311bda46f5c08386ead0d3a

Observation 06c7343a-0aec-4818-8248-c1f59cbebcfd · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning via Implicit Imitation Guidance Offline Reinforcement Learning with Implicit Q-Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.336559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.336559Z digest=sha256:794d7a19f0ccc831669399759879c6dcbe69ad00654ecb998d061215fbff1b41

Observation b2329927-ac49-4b9c-b574-274bd0fd61ed · outbound

This paper cites Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias.

Reinforcement Learning via Implicit Imitation Guidance Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.344696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.344696Z digest=sha256:cfdd87250949f4b1d1f7540027b78e652a7adb7c7593a084f56a056e93aee39f

Observation a13b6330-ea3a-4e6e-8793-87762ccd29da · outbound

This paper cites Over- coming exploration in reinforcement learning with demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Over- coming exploration in reinforcement learning with demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.349077Z digest=sha256:37360b188cd3ebd2f76d3856a01f4761861c2cf8346d30e39dec28a3899c1821

Observation 2d76a449-f55e-45b6-9a25-4fe50f74881b · outbound

This paper cites Computational Theories of Curiosity-Driven Learning.

Reinforcement Learning via Implicit Imitation Guidance Computational Theories of Curiosity-Driven Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.541101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.356660Z digest=sha256:a2186d4c66db7eed42b4d682ffe27b38b8310f40e00c25f733ddfe54ac5ca9c7

Observation aeaa12dd-e2cd-41c4-8703-41c978fb7710 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.368469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.368469Z digest=sha256:1c72dcb367f01732c3c16964e7268bb81ddfc93f9efb2acb097f9824b28de9b0

Observation 8a60e0dd-1666-43bd-8c84-0e5cc8d313af · outbound

This paper cites Schmidhuber.

Reinforcement Learning via Implicit Imitation Guidance Schmidhuber

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.661903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.376056Z digest=sha256:0a6d4341b6b903ce74c088f81f6e5da8da4f8820528ad21a3b55fae71c5941a4

Observation 13036d68-8079-4468-a573-ac508a627241 · outbound

This paper cites Parrot: Data-Driven Behavioral Priors for Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Parrot: Data-Driven Behavioral Priors for Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.379573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.379573Z digest=sha256:e6b844b22d18ba6cdc504e90eefae3a75acd0facd42a3306725a015b7b838797

Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.383535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.383535Z digest=sha256:6a46e792c7a72972de9118c7033e3b26ea3b6a3cb49ba0d8128ed42eeb17c1b6

Observation 4474229d-0083-4b59-87ee-9d01368b8b76 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Reinforcement Learning via Implicit Imitation Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.387296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.387296Z digest=sha256:fefa6cf8a28109bfca9f5cea1cdfd7fa25ef8b9c0f9eb299aacc8ec269881008

Observation 0066d2d0-370f-462f-b277-9688d2779557 · outbound

This paper cites Learning latent state representation for speeding up exploration.

Reinforcement Learning via Implicit Imitation Guidance Learning latent state representation for speeding up exploration

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.450428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.391312Z digest=sha256:5285104692f20c9d87722ae2189bb51c28432db24b6a3c402c9138dbe542ac27

Observation 5da8b341-7bea-4905-9769-d2ff68f8c0dd · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.395221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.395221Z digest=sha256:5239502081cbe1c3f4c39aae87a4943b347f64a6a4a182e55e91ab0f241a070e

Observation 14ab7d25-fb58-442a-b0b2-d46ed3aa6d56 · outbound

This paper cites Exploration by Random Network Distillation.

Reinforcement Learning via Implicit Imitation Guidance Exploration by Random Network Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.319807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.319807Z digest=sha256:138b6762051af54ea3758e46d46ee4458fe1a164ecb7455f7ca30f0ac6eaa5a7

Observation fd8c01a0-028f-4001-ae53-4883703f0cc9 · outbound

This paper cites Learning by Playing - Solving Sparse Reward Tasks from Scratch.

Reinforcement Learning via Implicit Imitation Guidance Learning by Playing - Solving Sparse Reward Tasks from Scratch

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.496733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:37:46.372311Z digest=sha256:2fc8fb30a9f4adc9c230f6a2cb5916179ece0e0c1bca86d0d62cd83a29e1866d

Observation 4197caaf-9d9d-41ea-b2dc-43b242038ebc · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Reinforcement Learning via Implicit Imitation Guidance AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.352606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.352606Z digest=sha256:7b7ceee48e315f546860e09155e1da0488b9de329ee4ba9ea26b49a51523dd67

Observation 084fe3ea-9951-45c1-9f97-44b59aa7afea · outbound

This paper cites Self-Supervised Exploration via Disagreement.

Reinforcement Learning via Implicit Imitation Guidance Self-Supervised Exploration via Disagreement

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.364579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.364579Z digest=sha256:538c8344a3f92ccad168eab1e1a556f1ba911f2a37ee9d8397791dc9db510a53

Observation 18e26509-53b6-4ac0-a579-ec7add5f2be0 · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Go-Explore: a New Approach for Hard-Exploration Problems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.323819Z digest=sha256:be85b1f87d4723eb2b0c4c50529b275989ca577168859c1a35b38558decb69a9

Observation 59835236-da50-444b-87cb-09f63dc28f21 · outbound

This paper cites Efficient Exploration via State Marginal Matching.

Reinforcement Learning via Implicit Imitation Guidance Efficient Exploration via State Marginal Matching

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.340669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.340669Z digest=sha256:f9bce5537f7d86952758980069127fd5a0c084eae6904118305366ead09c3d75

Observation 88953db2-4e1a-4aa0-8f4b-bd80c719fd01 · outbound

This paper cites Making Efficient Use of Demonstrations to Solve Hard Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Making Efficient Use of Demonstrations to Solve Hard Exploration Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.360602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.360602Z digest=sha256:2899656ba6882f78ab864e963f9bdf14a358d391e3336cd40ca217df3743d5e2

Observation 88e13b0c-c32f-4da1-9331-d149b31a9721 · outbound

This paper cites MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.328115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.328115Z digest=sha256:092532a3f00a360fe083c736c9e9f0953ee8d92bcfec28a42af1ce051a159aa6

Pith citing papers

Observation f7f7acb0-042e-4fab-b930-ab01711ff06d · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Reinforcement Learning via Implicit Imitation Guidance

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:05.252079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:8800676aaf3bc0771e20227aeefe1cb00da060a93d816e3d5ce0d6f783cf4a82

Observation 912f030f-4650-4d47-b137-548db7d98738 · inbound

Value Flows cites this paper.

Value Flows Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:27.929388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:27.929388Z digest=sha256:18e301950be1888c97e6b1d9b69a6667408453040bb422c38ac4618919e896b9

Observation 73bb73d3-8468-4d6c-b222-9bd2e736b647 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Reinforcement Learning via Implicit Imitation Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:59:59.505114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:fdd0e2cfa6fd8beca8f85f3a39ef9d96850483459ecd388fbd309aa6741b7fb5

Observation 740ec016-c10b-4d8c-b869-42282c0b9244 · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.992073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:48:36.372298Z digest=sha256:ae67a1b4c86b330b659c3c24fb262d7f7972feab80634e072a3e2034bbd251af

Observation 510bd843-417c-4782-93d9-602cb0adf3fb · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:08:53.253285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:08:53.253285Z digest=sha256:76ab51f129eff02080f9c3825640e4d91f1cf1dfd22646163bfe7699ecde9aee

Observation 2eee268b-bfe8-4404-b5b0-805b788a900a · inbound

FASTER: Value-Guided Sampling for Fast RL cites this paper.

FASTER: Value-Guided Sampling for Fast RL Reinforcement Learning via Implicit Imitation Guidance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.234174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:47:36.475845Z digest=sha256:35bc9646f2beef7c0bf2aeb2d77c925e9ab7dc313aa59034535416b6cb6c18d9

Observation 6dc67857-44d7-4bbf-85f2-b9e352933bc6 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.032376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.032376Z digest=sha256:b17507c1c256cbf90ca8bfa7b3a270c01ed8cf611b379768160a5cdf1d89b4b1

Observation 49b28c9e-8205-4ea0-b167-2c60e716a353 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.858412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.858412Z digest=sha256:d9a51da8739631e96fad8b7eba2f06c3f47fffd289ab7d151f0694d9d17c9474