Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:1812.02900.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:07.755855Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T12:35:48.973163Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 63b06953-ea16-4a6a-964a-e3a3380b921a · inbound
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog Off-Policy Deep Reinforcement Learning without Exploration
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6b042a1-08d5-4666-8bde-5a1126682347 · inbound
Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22f1d582-6358-4ddc-9f69-86b106be883b · inbound
D4RL: Datasets for Deep Data-Driven Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22bd0f83-d64b-421f-b16c-3c9cff3ce925 · inbound
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Off-Policy Deep Reinforcement Learning without Exploration
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cffc500b-6e73-442a-b4dd-fe35a2b93c91 · inbound
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields Off-Policy Deep Reinforcement Learning without Exploration
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a01e27ca-ae96-4cf7-8ea7-ad628b92dfc6 · inbound
A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Off-Policy Deep Reinforcement Learning without Exploration
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72888b8-ee2f-42b6-bb8d-62f8603480b2 · inbound
DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Off-Policy Deep Reinforcement Learning without Exploration
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59543e8c-047a-4967-8190-805e0c333707 · inbound
Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management Off-Policy Deep Reinforcement Learning without Exploration
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 205f6cf7-2b43-4d79-bc0c-0164ad81ea66 · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb0b81a7-af25-4529-8e26-ec16c1db2bcb · inbound
RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c6f62e2-ea4c-4def-abbd-e319c799d8e8 · inbound
Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Off-Policy Deep Reinforcement Learning without Exploration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6416329b-4cb6-43ff-b08c-718c17b387b5 · inbound
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · inbound
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20d9cb9-f75b-4cdc-9654-9a97e46cb8fc · inbound
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Off-Policy Deep Reinforcement Learning without Exploration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d55c0f0-601a-4975-94d9-2cc4805c1d1b · inbound
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Off-Policy Deep Reinforcement Learning without Exploration
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.