Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:33.558286Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2412.10778.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:33.558286Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ed2e75b-077b-4876-949a-a62e9bc3093b · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Human-level control through deep reinforcement learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f78713-3e70-4d4d-ad11-82bdfb58d759 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Continuous control with deep reinforcement learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5de010b6-9673-49dd-9151-831ae5e3ea94 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Balancing state exploration and skill diversity in unsupervised skill discovery,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ac017a37-d82b-47ef-92a8-4c2010a2d2ff · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos URLB: Unsupervised reinforcement learning benchmark,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c61a4e09-b102-4216-95f2-862dd6175a56 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Effective representation learning is more effective in reinforcement learning than you think,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99a9fef7-e0e0-4b70-aa28-ed4e90520bc3 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b3bdfb8-989a-4609-ba87-799b0c30edd5 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Human-level control through directly trained deep spiking q-networks,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cb5f13de-c83b-44ee-99f3-f9cdedcfb322 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Deep reinforcement learning-based automatic exploration for navigation in unknown environment,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f39e2bf-452a-46fe-bc25-7d06a25b6d90 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Model based reinforcement learning for atari,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b51610b5-ad8c-45ed-bbbb-7193d00240ae · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Dream to control: Learning behaviors by latent imagination,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7150e8d8-ec79-44fc-9112-9d6d96fe8901 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Prototypical context-aware dynamics for gener- alization in visual control with model-based reinforcement learning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b45358da-3250-4f90-a234-355855e0d2dd · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Mastering atari with discrete world models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b4e295-8987-4040-a883-3dccd86a2545 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Data-efficient reinforcement learning with self- predictive representations,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75c57e4c-0eab-42c1-9946-51aad1536342 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning with unsupervised auxiliary tasks,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 945fdef8-df89-4e98-b6ef-20ae0bf5b9ee · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked and inverse dynamics modeling for data-efficient reinforcement learning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2e7ae6f5-b4de-4792-aa5d-2619cbcb54d2 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning future representation with synthetic observations for sample-efficient reinforcement learning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c1c80603-0cf8-4ecf-afc9-1a341a610239 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Design from policies: Conservative test-time adaptation for offline policy optimization,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2ebd97ab-c429-4b4a-8061-1996a4e9d893 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Hiql: Offline goal-conditioned rl with latent states as actions,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dcd66665-e15e-4580-8347-2062c76fe5f2 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A survey of imitation learning: Algorithms, recent developments, and challenges,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a86abef-1e76-47f0-9d19-173f8d81fdcd · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Generative adversarial imitation learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 50fd8ef2-f333-43b3-8f10-d4f093d43eaf · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Robotic offline rl from inter- net videos via value-function learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b8cb51b-625e-4a30-9037-8501e69ed6a6 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning from passive data via latent intentions,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0bfce7e9-df19-42b7-9c4c-25b243d89072 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Diffusion reward: Learning rewards via conditional video diffusion,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a839e7bd-7fc0-4c02-aa5f-dad5ca4ca3a5 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Generative Adversarial Imitation from Observation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87404cf-9816-436c-83be-3c39d3fd6826 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning from visual observation via offline pretrained state-to-go transformer,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d4c5766-18cf-4014-8f9c-8721bb100b63 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Video prediction models as rewards for reinforcement learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1e5283bb-5ed2-4842-86c4-729eae8464fc · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Ilpo-mp: Mode priors prevent mode collapse when imitating latent policies from observations,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 017d5ac7-d16f-40da-aaa4-5a5907a83b75 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Behavioral cloning from obser- vation,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97d5b92-0e4b-4693-bdee-484552afcb7b · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Imitating latent policies from observation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a777e21-585c-44ee-a9de-6df24ef2f5d4 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Steps: Joint self-supervised nighttime image enhancement and depth estimation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb839d9c-41d2-4a90-aec5-54b49744cb68 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learn- ing with prototypical representations,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9ce33c-feff-4c41-a8f6-5b3773b56d2a · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Intrinsically motivated self- supervised learning in reinforcement learning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 23458c74-1c8a-4f5c-9e81-c362bf614def · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Deep reinforcement learning for autonomous driving: A survey,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab1d93a-620a-4ff1-bfd7-04c8e921c0a6 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Scalable deep reinforcement learning for vision-based robotic manipulation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aef02e9f-444a-4137-b4bb-4f2cc325b220 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Mastering Diverse Domains through World Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4436880e-2654-4eb8-b6ed-3ad56388f8e0 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representa- tions,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3356251f-21b9-46b5-82a0-1edae7ee10ea · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A survey on model-based reinforcement learning,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78ea1545-2bd2-4a86-afaa-0b9730b4e47a · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Curl: Contrastive unsupervised representations for reinforcement learning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83f10578-17bc-4f3f-af4b-2dad5d55dcf2 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked Visual Pre-training for Motor Control
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5242b6dd-189d-42cb-89f9-d7d51ba8f1e2 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Value-consistent representation learning for data-efficient reinforcement learning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20d17f4e-42d7-496a-b606-8ba878d5f664 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Cross-domain random pretraining with prototypes for reinforcement learning,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a8ee257-08a9-41b3-9b97-6738bc967c98 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Imitation learning: A survey of learning methods,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4208aab7-78ce-44ad-9058-3c828de85af9 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning with action-free pre-training from videos,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b65159b4-c32f-4350-9f56-7c3786898076 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Video pretraining (vpt): Learning to act by watching unlabeled online videos,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3fd6e46d-52bf-420a-9366-da5312ffe704 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked world models for visual control,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4455f16d-f677-4089-a0dc-94fa816436f1 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Multi-view masked world models for visual robotic manipulation,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fbb167bd-cebe-4eae-97ac-d4e9651ad1e5 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Visual imitation learning with patch rewards,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3aded90c-8958-436e-88ab-58d609be43a6 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Adversarial imitation learning from visual observations using latent information,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0cf2abea-33e2-44db-ba82-56717b014fbf · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Zero-shot visual imitation,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aaad3a52-59a3-4268-8275-abd662dc58f6 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Momentum contrast for unsupervised visual representation learning,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfeb7c60-bc68-4402-9cd8-94a2f6da509e · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Decoupling rep- resentation learning from reinforcement learning,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation feb31420-7e35-44d4-8d7a-ec93ac6b9f5d · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A simple frame- work for contrastive learning of visual representations,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12841362-dd65-4661-95be-6252a9f9d839 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos On mutual information maximization for representation learning,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1b9dc4f-4eb0-4fc2-bc0a-3b4554891fe1 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Neural discrete representa- tion learning,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3d819b2d-8b6e-476b-a2de-763ed1cf4834 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning to act without actions,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51af3637-8af6-41d7-90d3-d131e68ec2db · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Proximal Policy Optimization Algorithms
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59126c88-becc-4a92-b06c-98e23ad77629 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Policy gradient without boostrapping via truncated value learning,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c2b921c-6c76-4ab1-8575-b56f270a0911 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Leveraging procedu- ral generation to benchmark reinforcement learning,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 234d9e55-002e-4f43-85e9-5a0c83215da6 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14cdab2b-6972-4560-a26f-975e8a12bdf4 · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos U-net: Convolutional networks for biomedical image segmentation,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1898d2-b97c-4428-8317-491bc9b8b80d · outbound
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Adam: A Method for Stochastic Optimization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.