Pith. sign in

Paper Citation Record · LEDGER

Off-Policy Deep Reinforcement Learning without Exploration

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:1812.02900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1812.02900 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:49.667242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T12:35:48.973163Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63b06953-ea16-4a6a-964a-e3a3380b921a · inbound

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog cites this paper.

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog Off-Policy Deep Reinforcement Learning without Exploration

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T12:35:48.976218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T12:32:09.940698Z digest=sha256:c21e0e38dc2c082cbfef54230507c6a53a583f17abf143942adc1f5a4975a0bc

Observation b6b042a1-08d5-4666-8bde-5a1126682347 · inbound

Behavior Regularized Offline Reinforcement Learning cites this paper.

Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.423129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:0b59d6bbff590bb720aad175c7cdf1fbdaf73483db44d11e20e9f69675201255

Observation 22f1d582-6358-4ddc-9f69-86b106be883b · inbound

D4RL: Datasets for Deep Data-Driven Reinforcement Learning cites this paper.

D4RL: Datasets for Deep Data-Driven Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:19:17.409114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T23:19:17.322890Z digest=sha256:23ed5d7f1ce501813b47d1bc92d0f409f4f25fbfd864d6724cd15fc7e0cadea4

Observation 22bd0f83-d64b-421f-b16c-3c9cff3ce925 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Off-Policy Deep Reinforcement Learning without Exploration

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:33:21.574462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:411f99c731ba8a83422c2ef54123669387356df57e0e3248552c5ecebe83a8f9

Observation cffc500b-6e73-442a-b4dd-fe35a2b93c91 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields Off-Policy Deep Reinforcement Learning without Exploration

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:52:43.880492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:6201f74d231604f909fd3d7c52e4bcbbacc290f1de871849aa7402c094f63b3b

Observation 832ce9c9-889f-4ad2-9cb1-0298b845ddf0 · inbound

FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning cites this paper.

FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:49.667242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:49.667242Z digest=sha256:a659219ca8230678298c5b7e374f1b20867c307b19d8b8fb74a02aa15b84de85

Observation a01e27ca-ae96-4cf7-8ea7-ad628b92dfc6 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Off-Policy Deep Reinforcement Learning without Exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.755855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.755855Z digest=sha256:da5ce7a9da29edd5dac8813675d70e3835df925fcf5a72be4f217ebe292a6fef

Observation d72888b8-ee2f-42b6-bb8d-62f8603480b2 · inbound

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions cites this paper.

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Off-Policy Deep Reinforcement Learning without Exploration

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:56:26.092270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:55:16.513839Z digest=sha256:534dc17ba834bbef8d2139d86b1753e3852a078cdaa9bd01217a1b98aebab09a

Observation 59543e8c-047a-4967-8190-805e0c333707 · inbound

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management cites this paper.

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management Off-Policy Deep Reinforcement Learning without Exploration

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:44:14.995302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:40:31.084594Z digest=sha256:ebab78dba1ebe0b9500606de34d97ca3776ca27fd938670eb41719f8ace9b01b

Observation 205f6cf7-2b43-4d79-bc0c-0164ad81ea66 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.190747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:d929db2e2d950863f221e99b07e16130b5f0944ddca38800a1ee21bc6ea98f13

Observation fb0b81a7-af25-4529-8e26-ec16c1db2bcb · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.798839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:11bfb48ec7c1f123e967d14724d517ee70aaf1d23e7d7aa9cabe87c1e13d25cf

Observation 2c6f62e2-ea4c-4def-abbd-e319c799d8e8 · inbound

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation cites this paper.

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Off-Policy Deep Reinforcement Learning without Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T13:15:11.461088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:15:11.461088Z digest=sha256:baee6f85f5385308141b4a81232bba2e076a1cc7b012c58a8adf1acb9bbc466c

Observation 6416329b-4cb6-43ff-b08c-718c17b387b5 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.520577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.520577Z digest=sha256:03944d807949e03f58903ffd1a9aeb7f1727198f61e482014a023f63773b7590

Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.878960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.878960Z digest=sha256:127f0c38bea5ca1b768029414824455c9ba8aa6eb00b4339c936ece16f2c94ca

Observation e20d9cb9-f75b-4cdc-9654-9a97e46cb8fc · inbound

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search cites this paper.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Off-Policy Deep Reinforcement Learning without Exploration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T01:24:20.486289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:24:20.486289Z digest=sha256:a05c5f4ba6b2275a34832a496743271f411284683af392f774e371f9cf0d8d38

Observation 9d55c0f0-601a-4975-94d9-2cc4805c1d1b · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Off-Policy Deep Reinforcement Learning without Exploration

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:30.241728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:30.241728Z digest=sha256:97ee314ce310c87bf2027e39a48ee8e02dc8b64a69cacfc21e4d2d99e8343601