Pith. sign in

Paper Citation Record · LEDGER

Kickstarting Deep Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:1803.03835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1803.03835 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:09:27.590185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T21:47:36.272825Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bbcc9ef-b5f9-43d7-926e-fff712bd30c2 · inbound

Attentive Multi-Task Deep Reinforcement Learning cites this paper.

Attentive Multi-Task Deep Reinforcement Learning Kickstarting Deep Reinforcement Learning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T02:25:14.673926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T02:25:00.765509Z digest=sha256:f37ddecc31aedff25e59272d1a039649b3015d02f07fe54c100beff64fd241b8

Observation b06cf659-0c8d-4075-8eee-f93eac5afe06 · inbound

Red Teaming Language Models with Language Models cites this paper.

Red Teaming Language Models with Language Models Kickstarting Deep Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:57:51.836919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T20:57:51.360357Z digest=sha256:e91e44f00f5adf2555431265a87e55e1a607f2a7bb4105d49e0f5bc3aff4b6c3

Observation b5f2b6c2-94f6-4b43-89a0-7f17f3947cb0 · inbound

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models cites this paper.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Kickstarting Deep Reinforcement Learning

Reference 287

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:43:30.299758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:79ec3ede02c8faf1fbb6b47feb39bb931232c4422e2f492cf8b2c1bd0bf455e0

Observation 0070de22-2eaa-4108-927a-28b1890f182b · inbound

Proximal Policy Distillation cites this paper.

Proximal Policy Distillation Kickstarting Deep Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:13:36.922605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T23:09:42.421333Z digest=sha256:04b33e02597d1d9dfa9a2e163c1a2f969c316c856560ffcfe7fc4e4cd1044d49

Observation b3d45bd2-b340-4d88-ae73-794929bd22c4 · inbound

Towards an Autonomous Test Driver: High-Performance Driver Modeling via Reinforcement Learning cites this paper.

Towards an Autonomous Test Driver: High-Performance Driver Modeling via Reinforcement Learning Kickstarting Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:09:27.590185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:09:27.590185Z digest=sha256:d6b1da018257e95f4deafc7bf7582c17b7aea22415f58b89aca8ac7e52cf6f1f

Observation 92c2c69f-d3e7-4a1b-bb6f-741fa4ca8df8 · inbound

All You Need in Knowledge Distillation Is a Tailored Coordinate System cites this paper.

All You Need in Knowledge Distillation Is a Tailored Coordinate System Kickstarting Deep Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:11:13.064499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:11:13.064499Z digest=sha256:bea32c5de2880f34adec8ca850518a6c8e17b24760ee14be455a2e9abe495c6f

Observation 0b5891fb-3032-4457-a142-a333a4535392 · inbound

Embodied CoT Distillation From LLM To Off-the-shelf Agents cites this paper.

Embodied CoT Distillation From LLM To Off-the-shelf Agents Kickstarting Deep Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:14.680961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:14.680961Z digest=sha256:ceeba1f5c5a8d6dfab080fdda5d9a61a33536864422afc94be09367868c29b1d

Observation 0c28a6a4-bf45-4af6-8a3d-eac8f7dd2c37 · inbound

Energy-Based Transfer for Reinforcement Learning cites this paper.

Energy-Based Transfer for Reinforcement Learning Kickstarting Deep Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.762144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.762144Z digest=sha256:d2d39418bf3e1f7ff8be3c8d05b74ba70af81c147fc270e48f3aa404431aa5c6

Observation 0b143ffd-91bf-4246-9cb8-0f381273ab27 · inbound

ThinkTuning: Instilling Cognitive Reflections without Distillation cites this paper.

ThinkTuning: Instilling Cognitive Reflections without Distillation Kickstarting Deep Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T22:05:05.921845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:05:05.921845Z digest=sha256:c5d2cc3a862c6365ca39b1e7efa28e06c66722ef91a0b0b39afe2438c84b0970

Observation 1f63d06b-0b17-4988-93e5-828fe488091a · inbound

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance cites this paper.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Kickstarting Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.083396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.083396Z digest=sha256:90e5e36a823d05bf9fa553116a5234eb45d2fccd71c67eedbba9c96204ae8310

Observation 63603a1f-eb18-448c-9a70-5af36f61f324 · inbound

TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations cites this paper.

TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations Kickstarting Deep Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:38:49.978324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T02:27:51.946679Z digest=sha256:5fb9f47a3897a8439ba99f85d1fec9679e0f206c8d1b388eeca8f727bea24cff

Observation 665fdad2-4582-4307-9275-b3d615d47f89 · inbound

TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations cites this paper.

TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations Kickstarting Deep Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T11:08:39.028689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:08:39.028689Z digest=sha256:cb27e853cbbc1db0208912fadd9d86dd14ba89234ebb3f8e19367799e52304f9

Observation d55c6583-052f-4f63-8556-312c744e48f4 · inbound

MotionPyramid: Hierarchical Motion Representation and Residual Interfaces cites this paper.

MotionPyramid: Hierarchical Motion Representation and Residual Interfaces Kickstarting Deep Reinforcement Learning

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:44.290147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T04:14:41.060330Z digest=sha256:3cc4b28da4f2450d7a3f28d93e54696bb7f21aabbab5555710b7e3418b984407

Observation 52ca7b67-a72c-4c0f-b474-31b7fc32263c · inbound

Efficient Long-Horizon Learning for Learned Optimization cites this paper.

Efficient Long-Horizon Learning for Learned Optimization Kickstarting Deep Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-10T21:47:36.289563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T21:40:06.294302Z digest=sha256:2d8052faf09aad00770c8e785135bf7e79c39eec7c3960b2b085cf14a8e504c8

Observation d380fda5-b604-4a91-8d03-f0e006fcaa4a · inbound

Efficient Long-Horizon Learning for Learned Optimization cites this paper.

Efficient Long-Horizon Learning for Learned Optimization Kickstarting Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T15:58:06.643463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:58:06.643463Z digest=sha256:73e8ea019d123f6ce5a20dc5032f48d16ce93b1455847ecdd4ee325788d8632d

Observation f6b3ff2b-e8ba-4721-bba3-699c5eb6393d · inbound

Efficient Long-Horizon Learning for Learned Optimization cites this paper.

Efficient Long-Horizon Learning for Learned Optimization Kickstarting Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T08:16:52.647887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:16:52.647887Z digest=sha256:53743238cf8925c143fc1ebdb0cd680f82cc0924a55d86b74fe6ce2873604321

Observation d68ed56d-7968-4fa7-afb0-f96d69f4236a · inbound

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models cites this paper.

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models Kickstarting Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:39:40.941485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:39:40.941485Z digest=sha256:0d014859451da14b6a38c5341b0658139401b003dd584ca1e5b39d86d9889113