Pith. sign in

Paper Citation Record · LEDGER

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos

As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2412.10778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10778 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:33.558286Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ed2e75b-077b-4876-949a-a62e9bc3093b · outbound

This paper cites Human-level control through deep reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.224078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.224078Z digest=sha256:b3622d0830246dff7294f3c253de1284c93a1758b1baba6d3c22152d248343c1

Observation 75f78713-3e70-4d4d-ad11-82bdfb58d759 · outbound

This paper cites Continuous control with deep reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Continuous control with deep reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.796402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.230030Z digest=sha256:f0eec9d6cd804146d75c6fbe43262ab60f1c2ebc4c6518734605b87c86ab9dfc

Observation 5de010b6-9673-49dd-9151-831ae5e3ea94 · outbound

This paper cites Balancing state exploration and skill diversity in unsupervised skill discovery,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Balancing state exploration and skill diversity in unsupervised skill discovery,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.778580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.235252Z digest=sha256:2d8ae98bc810da0a4ae12cc6224f75178235642a5cb610a51c1626794837529d

Observation ac017a37-d82b-47ef-92a8-4c2010a2d2ff · outbound

This paper cites URLB: Unsupervised reinforcement learning benchmark,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos URLB: Unsupervised reinforcement learning benchmark,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.758897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.240468Z digest=sha256:b3417139bb8b735c0fa7606382a67d236bbeb59aac1bb3325a8802de9a379031

Observation c61a4e09-b102-4216-95f2-862dd6175a56 · outbound

This paper cites Effective representation learning is more effective in reinforcement learning than you think,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Effective representation learning is more effective in reinforcement learning than you think,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.738619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.245565Z digest=sha256:619e3e7871001ca76cf06b6ecaf858e30637758ff702e4b71c1dccc7f60bd8c0

Observation 99a9fef7-e0e0-4b70-aa28-ed4e90520bc3 · outbound

This paper cites Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.721312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.250551Z digest=sha256:d5b1cf1fd868e7029152951a8d41788d9ab5a2e1e628080d693f0c5b44148adc

Observation 2b3bdfb8-989a-4609-ba87-799b0c30edd5 · outbound

This paper cites Human-level control through directly trained deep spiking q-networks,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Human-level control through directly trained deep spiking q-networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.703923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.256568Z digest=sha256:f50474104987fa209573316832e0f66727f4f29f3b4aa00fbad5d0ca0f0e5673

Observation cb5f13de-c83b-44ee-99f3-f9cdedcfb322 · outbound

This paper cites Deep reinforcement learning-based automatic exploration for navigation in unknown environment,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Deep reinforcement learning-based automatic exploration for navigation in unknown environment,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.684276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.261642Z digest=sha256:368c4fedd9773b4f52bec0192f5fcc458206b1664a2734d0630f27f51dfb5f18

Observation 0f39e2bf-452a-46fe-bc25-7d06a25b6d90 · outbound

This paper cites Model based reinforcement learning for atari,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Model based reinforcement learning for atari,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.663078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.266758Z digest=sha256:460c55b08c5f2947b49e7f5efbe75880ae72f17612cad0ff53ca662e27a3d27d

Observation b51610b5-ad8c-45ed-bbbb-7193d00240ae · outbound

This paper cites Dream to control: Learning behaviors by latent imagination,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Dream to control: Learning behaviors by latent imagination,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.271917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.271917Z digest=sha256:e05b842da0ce67b75e3ebd6c7a0ff065cde7d7dcc22c581fb6554b14a7823a14

Observation 7150e8d8-ec79-44fc-9112-9d6d96fe8901 · outbound

This paper cites Prototypical context-aware dynamics for gener- alization in visual control with model-based reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Prototypical context-aware dynamics for gener- alization in visual control with model-based reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.631752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.277054Z digest=sha256:2aa8020285b041eaf24af34ee3f23efeebb22e53e00ee9b1cb5957c6e6d32ac4

Observation b45358da-3250-4f90-a234-355855e0d2dd · outbound

This paper cites Mastering atari with discrete world models,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Mastering atari with discrete world models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.283182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.283182Z digest=sha256:123fd6d2f64169c69778e31cfa64a64e9ffc5da54e24d8ec01283af2c753ad7b

Observation 72b4e295-8987-4040-a883-3dccd86a2545 · outbound

This paper cites Data-efficient reinforcement learning with self- predictive representations,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Data-efficient reinforcement learning with self- predictive representations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.600499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.289176Z digest=sha256:c8189f6b5f143ebbc51b0048f62c9942114b47615e62271766123b0ab4db4b31

Observation 75c57e4c-0eab-42c1-9946-51aad1536342 · outbound

This paper cites Reinforcement learning with unsupervised auxiliary tasks,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning with unsupervised auxiliary tasks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.580243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.294585Z digest=sha256:845fb9ce222a537ff982ea95402156b81115f3948764fb5d2e89cdbfdcc915c7

Observation 945fdef8-df89-4e98-b6ef-20ae0bf5b9ee · outbound

This paper cites Masked and inverse dynamics modeling for data-efficient reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked and inverse dynamics modeling for data-efficient reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.563656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.300543Z digest=sha256:c887eb3b26aeff90938b45b3bc392ceb5a4a4956ba842cd87ad8b57110f5a93f

Observation 2e7ae6f5-b4de-4792-aa5d-2619cbcb54d2 · outbound

This paper cites Learning future representation with synthetic observations for sample-efficient reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning future representation with synthetic observations for sample-efficient reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.545905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.306734Z digest=sha256:b6ee560aaf896d031ebfda0b0355bbb292eac7571a68c31a5b6f58cb1ced16fe

Observation c1c80603-0cf8-4ecf-afc9-1a341a610239 · outbound

This paper cites Design from policies: Conservative test-time adaptation for offline policy optimization,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Design from policies: Conservative test-time adaptation for offline policy optimization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.529294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.312306Z digest=sha256:e3de281fc39bcad5923107b2aedc1a5a2175f51779640831b11f49516149641b

Observation 2ebd97ab-c429-4b4a-8061-1996a4e9d893 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Hiql: Offline goal-conditioned rl with latent states as actions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.513012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.318820Z digest=sha256:6e6fd9e7c6134e0bf28cc63babbe9b31cf166ca260aa105a6f970d4a4f15075c

Observation dcd66665-e15e-4580-8347-2062c76fe5f2 · outbound

This paper cites A survey of imitation learning: Algorithms, recent developments, and challenges,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A survey of imitation learning: Algorithms, recent developments, and challenges,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.324977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.324977Z digest=sha256:646975abf927743fc4d47e03816c4e97f113aca26533d65c4f8aae7af9b78926

Observation 7a86abef-1e76-47f0-9d19-173f8d81fdcd · outbound

This paper cites Generative adversarial imitation learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Generative adversarial imitation learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.483974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.330351Z digest=sha256:185a28e551821d728fc6cb48459f29f483e7da5312b82097d929c076610b22cd

Observation 50fd8ef2-f333-43b3-8f10-d4f093d43eaf · outbound

This paper cites Robotic offline rl from inter- net videos via value-function learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Robotic offline rl from inter- net videos via value-function learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.464282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.335615Z digest=sha256:65a76cb2c74cca2fe652d4c4d2d8dbe5bf3d731df4ec1d954f19e6a75b6ca7f8

Observation 2b8cb51b-625e-4a30-9037-8501e69ed6a6 · outbound

This paper cites Reinforcement learning from passive data via latent intentions,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning from passive data via latent intentions,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.440902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.341267Z digest=sha256:6e053324478db8093bc1f0c484871305f9b76ba002f8a944c39190cb4142caa3

Observation 0bfce7e9-df19-42b7-9c4c-25b243d89072 · outbound

This paper cites Diffusion reward: Learning rewards via conditional video diffusion,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Diffusion reward: Learning rewards via conditional video diffusion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.421039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.347723Z digest=sha256:77ac21129eda58ea8da7093eaeb43bbdce4ba7ccf73cb877c324a1623db24ba1

Observation a839e7bd-7fc0-4c02-aa5f-dad5ca4ca3a5 · outbound

This paper cites Generative Adversarial Imitation from Observation.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Generative Adversarial Imitation from Observation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.353982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.353982Z digest=sha256:b08f8f6cd3851faf710f62ec90d71bce0a836b4e5a925de1eb50924a841bdc90

Observation d87404cf-9816-436c-83be-3c39d3fd6826 · outbound

This paper cites Learning from visual observation via offline pretrained state-to-go transformer,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning from visual observation via offline pretrained state-to-go transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.395230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.359366Z digest=sha256:4f4676b84b4d671b24621d172fbd9142702bed56ada29c97d970e692a3edc271

Observation 0d4c5766-18cf-4014-8f9c-8721bb100b63 · outbound

This paper cites Video prediction models as rewards for reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Video prediction models as rewards for reinforcement learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.375820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.364350Z digest=sha256:12ed6ccfd25cc0089224b4f1d405cd37a2aec1cc9e2858033f4c4f83efa5556c

Observation 1e5283bb-5ed2-4842-86c4-729eae8464fc · outbound

This paper cites Ilpo-mp: Mode priors prevent mode collapse when imitating latent policies from observations,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Ilpo-mp: Mode priors prevent mode collapse when imitating latent policies from observations,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.351401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.369425Z digest=sha256:31a4decf63e31777324a9f610e95f8540b4c160b55a83a183a17231f8bb362e9

Observation 017d5ac7-d16f-40da-aaa4-5a5907a83b75 · outbound

This paper cites Behavioral cloning from obser- vation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Behavioral cloning from obser- vation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.374553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.374553Z digest=sha256:26ae6acc94e5a625ed39cebeef07699c04b2e41d6a90cd45ccc9b3925570780e

Observation d97d5b92-0e4b-4693-bdee-484552afcb7b · outbound

This paper cites Imitating latent policies from observation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Imitating latent policies from observation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.313740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.379850Z digest=sha256:bb7fdf09e9d99da2b5851c4b4567a764057f63729cd0c2618fec0d7cfff55494

Observation 2a777e21-585c-44ee-a9de-6df24ef2f5d4 · outbound

This paper cites Steps: Joint self-supervised nighttime image enhancement and depth estimation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Steps: Joint self-supervised nighttime image enhancement and depth estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.288541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.386170Z digest=sha256:d0071460f88b7ce8c1dc4a10ca7b8b4b571f7fc75fd0b0aad3511c783cffd63f

Observation eb839d9c-41d2-4a90-aec5-54b49744cb68 · outbound

This paper cites Reinforcement learn- ing with prototypical representations,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learn- ing with prototypical representations,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.392657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.392657Z digest=sha256:1419ffe35907369c3017db76dba4de14506d80762abbe8f98b5970ae310ad8a4

Observation ea9ce33c-feff-4c41-a8f6-5b3773b56d2a · outbound

This paper cites Intrinsically motivated self- supervised learning in reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Intrinsically motivated self- supervised learning in reinforcement learning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.251760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.398455Z digest=sha256:4887acf7a234bd4ddb9e8f033d8aabac107ead18856d2f4c5aa18f7b466b0917

Observation 23458c74-1c8a-4f5c-9e81-c362bf614def · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Deep reinforcement learning for autonomous driving: A survey,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.403630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.403630Z digest=sha256:9d860c4edef74908a8b2445546844c2130b43561ffd7e6ba77b06e6cad74e224

Observation 5ab1d93a-620a-4ff1-bfd7-04c8e921c0a6 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Scalable deep reinforcement learning for vision-based robotic manipulation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.216852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.409028Z digest=sha256:02b56180b461fc4025cfe2150cb5a3215ecb25e67ac17cb840ecb02df10edb33

Observation aef02e9f-444a-4137-b4bb-4f2cc325b220 · outbound

This paper cites Mastering Diverse Domains through World Models.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Mastering Diverse Domains through World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.413904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.413904Z digest=sha256:d6944d2bc731eed6c503bbd20e0782f68bec96bd6cdda74beae7ec321f5cf19f

Observation 4436880e-2654-4eb8-b6ed-3ad56388f8e0 · outbound

This paper cites Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representa- tions,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representa- tions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.180883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.419892Z digest=sha256:c5556514f4439a5ff4b10942b601a408260514bd616249b0c6e7c16b67a637c9

Observation 3356251f-21b9-46b5-82a0-1edae7ee10ea · outbound

This paper cites A survey on model-based reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A survey on model-based reinforcement learning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.158970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.425363Z digest=sha256:e3be25a74587741f3f5d2ce6541a4f46f44c86e0f6a4b870c73372df51dfb26c

Observation 78ea1545-2bd2-4a86-afaa-0b9730b4e47a · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Curl: Contrastive unsupervised representations for reinforcement learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.140516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.430342Z digest=sha256:5a461e2bd8cc75616662d587fc809ede9bb64e4693c9b20fda8a14edfc45e72f

Observation 83f10578-17bc-4f3f-af4b-2dad5d55dcf2 · outbound

This paper cites Masked Visual Pre-training for Motor Control.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked Visual Pre-training for Motor Control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.435199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.435199Z digest=sha256:3299b34b03aa6a145639a2c26963c5c0df28ce150e7b3761d12a0a1a16511260

Observation 5242b6dd-189d-42cb-89f9-d7d51ba8f1e2 · outbound

This paper cites Value-consistent representation learning for data-efficient reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Value-consistent representation learning for data-efficient reinforcement learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.123274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.440609Z digest=sha256:80d2e5fabc2b9d9635ccbc2124ddb1ac6651180fd77fcac620fada82b1d940af

Observation 20d17f4e-42d7-496a-b606-8ba878d5f664 · outbound

This paper cites Cross-domain random pretraining with prototypes for reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Cross-domain random pretraining with prototypes for reinforcement learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.104164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.445845Z digest=sha256:4208b3b9cfa91dfc0d6704f7abc615d1dace8e98340417bd082eee38a9382cbf

Observation 2a8ee257-08a9-41b3-9b97-6738bc967c98 · outbound

This paper cites Imitation learning: A survey of learning methods,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Imitation learning: A survey of learning methods,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.451288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.451288Z digest=sha256:dd6c75a89e195cfd1c96b302978345e148b065bb8940e205921ea31c9e5f2dba

Observation 4208aab7-78ce-44ad-9058-3c828de85af9 · outbound

This paper cites Reinforcement learning with action-free pre-training from videos,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Reinforcement learning with action-free pre-training from videos,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.073417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.456781Z digest=sha256:0866d289032cb829ff4f2574ff3eb8ea69d9cb622770e894e4344f20b68be640

Observation b65159b4-c32f-4350-9f56-7c3786898076 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Video pretraining (vpt): Learning to act by watching unlabeled online videos,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.053724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.461587Z digest=sha256:555a3172bdaa377b6a59c1710f38403ec6322e3c33ab772205d54eccbd9368a2

Observation 3fd6e46d-52bf-420a-9366-da5312ffe704 · outbound

This paper cites Masked world models for visual control,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Masked world models for visual control,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.466888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.466888Z digest=sha256:1afd6b1859441eeaafef52459312a37631a1d4ceefefd131127a42c3c5afcb1c

Observation 4455f16d-f677-4089-a0dc-94fa816436f1 · outbound

This paper cites Multi-view masked world models for visual robotic manipulation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Multi-view masked world models for visual robotic manipulation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.022166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.472266Z digest=sha256:c43de9a9dbc35aefbeb7b208e01e29e82a6cab544049f57ae937ef7c02657145

Observation fbb167bd-cebe-4eae-97ac-d4e9651ad1e5 · outbound

This paper cites Visual imitation learning with patch rewards,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Visual imitation learning with patch rewards,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:34.000764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.477600Z digest=sha256:ef027981581f587a18ccbbcae4bacf03607852c34a26ea60bff38d99067f4a99

Observation 3aded90c-8958-436e-88ab-58d609be43a6 · outbound

This paper cites Adversarial imitation learning from visual observations using latent information,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Adversarial imitation learning from visual observations using latent information,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.978096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.483782Z digest=sha256:de65999b2ac9842935081882d7bcfa899d0fb5d0844a4e5f2a609ea6652d1ff6

Observation 0cf2abea-33e2-44db-ba82-56717b014fbf · outbound

This paper cites Zero-shot visual imitation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Zero-shot visual imitation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.956052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.488727Z digest=sha256:01c853da0bd67c5f001e519c0b2450f54d4c65b7eb8bdec246a1591edd43aeef

Observation aaad3a52-59a3-4268-8275-abd662dc58f6 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Momentum contrast for unsupervised visual representation learning,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.493603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.493603Z digest=sha256:8958edcca3cc4940f655378dba06fdf7dfaaa2a03cc0eb524af01edf223aa911

Observation bfeb7c60-bc68-4402-9cd8-94a2f6da509e · outbound

This paper cites Decoupling rep- resentation learning from reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Decoupling rep- resentation learning from reinforcement learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.920941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.499464Z digest=sha256:ac0f9d206073ed0a7ece1527b3aad1c00bcc25c94c5e7b199beba278448b6204

Observation feb31420-7e35-44d4-8d7a-ec93ac6b9f5d · outbound

This paper cites A simple frame- work for contrastive learning of visual representations,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos A simple frame- work for contrastive learning of visual representations,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.901313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.505537Z digest=sha256:1e03fe1da7dafd7c88820e2ebb23e872b93d8e6c98eb7df8eb4d26f3975364c2

Observation 12841362-dd65-4661-95be-6252a9f9d839 · outbound

This paper cites On mutual information maximization for representation learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos On mutual information maximization for representation learning,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.879892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.511798Z digest=sha256:b597e3ebb1127ce57df98d564c63cb64c32173246c10a96d75b1f66d577a62b6

Observation e1b9dc4f-4eb0-4fc2-bc0a-3b4554891fe1 · outbound

This paper cites Neural discrete representa- tion learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Neural discrete representa- tion learning,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.849930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.517798Z digest=sha256:2afd03b2e8f578564717eb81d7b2a7b7845dea3c615a973dd98faa0acd6bf3d0

Observation 3d819b2d-8b6e-476b-a2de-763ed1cf4834 · outbound

This paper cites Learning to act without actions,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Learning to act without actions,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.819956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.523928Z digest=sha256:97893545b0bf60544ba50a3e03f9534f55e59e181e7ce4f7b39b08ceff129cf8

Observation 51af3637-8af6-41d7-90d3-d131e68ec2db · outbound

This paper cites Proximal Policy Optimization Algorithms.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Proximal Policy Optimization Algorithms

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.528953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.528953Z digest=sha256:3280c75c423202f9d1c86cf9e1fdf44ce1b057613b0ab5e3ff6eb7a17ea307bc

Observation 59126c88-becc-4a92-b06c-98e23ad77629 · outbound

This paper cites Policy gradient without boostrapping via truncated value learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Policy gradient without boostrapping via truncated value learning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.795563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.534314Z digest=sha256:7e8255cddb1ab515b4d4900fe2a017e7f3e2ea42989df1909e44369b22a6d5a9

Observation 4c2b921c-6c76-4ab1-8575-b56f270a0911 · outbound

This paper cites Leveraging procedu- ral generation to benchmark reinforcement learning,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Leveraging procedu- ral generation to benchmark reinforcement learning,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.767644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.539412Z digest=sha256:afaed9e697b9b33189d0cf56e352154a0985f56d4a7425cf6a8f5789d0a4d6ec

Observation 234d9e55-002e-4f43-85e9-5a0c83215da6 · outbound

This paper cites Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:33.745728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:42:33.547687Z digest=sha256:9dff7812a8f88c30e0383b8181d64c3eb56816eb98539f55ed48be5980b769fc

Observation 14cdab2b-6972-4560-a26f-975e8a12bdf4 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos U-net: Convolutional networks for biomedical image segmentation,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.553005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.553005Z digest=sha256:ad4f70b93e24f2d613cb059e96e7654ce67c7238b6da618384307af61c8fb5b8

Observation bc1898d2-b97c-4428-8317-491bc9b8b80d · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos Adam: A Method for Stochastic Optimization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:33.558286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:33.558286Z digest=sha256:ddc2c6be60837605fcdd5ebe7a10d4f4b9591c9a4c4f6ebeced09318e86505eb

Pith citing papers

No inbound Pith citation observations are available.