Pith. sign in

Paper Citation Record · LEDGER

Off-Policy Deep Reinforcement Learning without Exploration

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:1812.02900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1812.02900 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:07.755855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T12:35:48.973163Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63b06953-ea16-4a6a-964a-e3a3380b921a · inbound

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog cites this paper.

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog Off-Policy Deep Reinforcement Learning without Exploration

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T12:35:48.976218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T12:32:09.940698Z digest=sha256:a5d3e744f9b064963eada79a08944105b46abb1dc0130e2aa6a053c55dde01a8

Observation b6b042a1-08d5-4666-8bde-5a1126682347 · inbound

Behavior Regularized Offline Reinforcement Learning cites this paper.

Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.423129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:899e1151d09d5058fe2ed1cfc73f683ce5e632f9b200509ddc5199ef46168c94

Observation 22f1d582-6358-4ddc-9f69-86b106be883b · inbound

D4RL: Datasets for Deep Data-Driven Reinforcement Learning cites this paper.

D4RL: Datasets for Deep Data-Driven Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:19:17.409114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T23:19:17.322890Z digest=sha256:cc8b59593f8c94d79e875943a8a25238ca40c82c6543c73ed902295213d45fa6

Observation 22bd0f83-d64b-421f-b16c-3c9cff3ce925 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Off-Policy Deep Reinforcement Learning without Exploration

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:33:21.574462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:81daa0e6ba5fbe6f2880cca4c3b1858d29ac82f298b2a6f1c18c678c85c3f878

Observation cffc500b-6e73-442a-b4dd-fe35a2b93c91 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields Off-Policy Deep Reinforcement Learning without Exploration

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:52:43.880492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:55585c5d99fe59da3a35f2b57fc6354aba341f9ba4d1a547a8447a205d2cc2fe

Observation a01e27ca-ae96-4cf7-8ea7-ad628b92dfc6 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Off-Policy Deep Reinforcement Learning without Exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.755855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.755855Z digest=sha256:da5ce7a9da29edd5dac8813675d70e3835df925fcf5a72be4f217ebe292a6fef

Observation d72888b8-ee2f-42b6-bb8d-62f8603480b2 · inbound

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions cites this paper.

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Off-Policy Deep Reinforcement Learning without Exploration

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:56:26.092270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:55:16.513839Z digest=sha256:b4f494ef7bdcccae6a34c1555d6840d7dd8c3fba42842b1a9818bc106cacfaa3

Observation 59543e8c-047a-4967-8190-805e0c333707 · inbound

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management cites this paper.

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management Off-Policy Deep Reinforcement Learning without Exploration

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:44:14.995302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:40:31.084594Z digest=sha256:d783a46c43150a0b2190baf132f6d8dad4ad427b77efb52d0043eae7501434cb

Observation 205f6cf7-2b43-4d79-bc0c-0164ad81ea66 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.190747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:dc215bb21da89e0c0b0a1c4f82e4d9db9c15995585cd25b2d7936ec1b119cdab

Observation fb0b81a7-af25-4529-8e26-ec16c1db2bcb · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.798839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:13199ee85c80af53564a71a3b6b45621eacc5da85809e0245205939354ba2c7b

Observation 2c6f62e2-ea4c-4def-abbd-e319c799d8e8 · inbound

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation cites this paper.

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Off-Policy Deep Reinforcement Learning without Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T13:15:11.461088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:15:11.461088Z digest=sha256:4f080a17dc83402314d1e16eda9f5a6748afe4d4c56c23bfe3ee9613cff5b31a

Observation 6416329b-4cb6-43ff-b08c-718c17b387b5 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.520577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.520577Z digest=sha256:d2c1d08f448ee1b5d736ef722fb0263673eba337c46ca7c3ad46f5b319996d91

Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.878960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.878960Z digest=sha256:2a5058731ea24610033585c5483a9a1c0510e672fbcd4c0fa2c6863188dace6b

Observation e20d9cb9-f75b-4cdc-9654-9a97e46cb8fc · inbound

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search cites this paper.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Off-Policy Deep Reinforcement Learning without Exploration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T01:24:20.486289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:24:20.486289Z digest=sha256:5e945aab2400b95a0f4d180479bdcd00122a4eaa0667af4de89149814e36d033

Observation 9d55c0f0-601a-4975-94d9-2cc4805c1d1b · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Off-Policy Deep Reinforcement Learning without Exploration

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:30.241728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:30.241728Z digest=sha256:05c4f72963689a6141c09ef428b1c5a72090c61f2946024efd57d439ceb07f16