Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:21:23.104750Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2501.02330.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:21:23.104750Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b67c23e-9c47-49c7-b547-a4314e23899c · outbound
SR-Reward: Taking The Path More Traveled Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45d5f5d2-17b0-41d4-b5d2-1f7abf21edcc · outbound
SR-Reward: Taking The Path More Traveled Holo-Dex: Teaching Dexterity with Immersive Mixed Reality
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f446d9b-6532-41f4-8b6f-d4dfe2c7474c · outbound
SR-Reward: Taking The Path More Traveled Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b7a34cf-ab4a-4c77-94ff-d057352f0d79 · outbound
SR-Reward: Taking The Path More Traveled Successor features for transfer in reinforcement learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7c0a29f-cc80-4f15-871b-1ead3ef0b752 · outbound
SR-Reward: Taking The Path More Traveled Mankowitz, Hado van Hasselt, R \' e mi Munos, David Silver, and Tom Schaul
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53128db3-72b6-48c0-a928-6b2bbb6b5244 · outbound
SR-Reward: Taking The Path More Traveled Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb19ae1-8db5-4144-90dd-99d09d2097cf · outbound
SR-Reward: Taking The Path More Traveled Successor feature sets: Generalizing successor representations across policies
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a06bce04-d870-4e87-9cfb-a7701e64f92b · outbound
SR-Reward: Taking The Path More Traveled Deep reinforcement learning from human preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccafeb6d-ad73-414e-928c-1ca7945b7ecb · outbound
SR-Reward: Taking The Path More Traveled Improving generalization for temporal difference learning: The successor representation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6fef980f-accc-4aa1-85d1-54d43be59acb · outbound
SR-Reward: Taking The Path More Traveled Model alignment as prospect theoretic optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bb78bb06-bbcf-4b1d-b8dd-9e76232d41a9 · outbound
SR-Reward: Taking The Path More Traveled Psiphi-learning: Reinforcement learning with demonstrations using successor features and inverse temporal difference learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b84edfd-463f-457d-8b94-7a7ee6b92a03 · outbound
SR-Reward: Taking The Path More Traveled Learning robust rewards with adverserial inverse reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a38220d3-3d03-43c0-93af-1d0600590a04 · outbound
SR-Reward: Taking The Path More Traveled D4rl: Datasets for deep data-driven reinforcement learning, 2020
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b032dc46-b8d2-47ed-ac25-587239e7229f · outbound
SR-Reward: Taking The Path More Traveled Addressing function approximation error in actor-critic methods
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c21cba52-019b-4603-8f2b-5b543a75985c · outbound
SR-Reward: Taking The Path More Traveled Off-policy deep reinforcement learning without exploration
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0473402f-a42b-4a25-8cc5-9fb1d6d03623 · outbound
SR-Reward: Taking The Path More Traveled For sale: State-action representation learning for deep reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5b6a7b7-ba45-4839-b6a0-fbc48437ce4d · outbound
SR-Reward: Taking The Path More Traveled Iq-learn: Inverse soft-q learning for imitation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec67118d-7084-4e4e-9012-c5b25488bf0a · outbound
SR-Reward: Taking The Path More Traveled Extreme q-learning: Maxent RL without entropy
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e466d414-70cd-49ac-8deb-ff3198d53b83 · outbound
SR-Reward: Taking The Path More Traveled A divergence minimization perspective on imitation learning methods, 2019
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 46076099-306e-4647-a5f8-2df7d07e80b5 · outbound
SR-Reward: Taking The Path More Traveled Generative Adversarial Networks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4d58a3-5be3-460b-9b16-11712bd38b1f · outbound
SR-Reward: Taking The Path More Traveled Maniskill2: A unified benchmark for generalizable manipulation skills
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5125191d-55c1-4633-aa9a-5761f45d6a23 · outbound
SR-Reward: Taking The Path More Traveled Generative adversarial imitation learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 824744c3-32a9-4835-a5e6-916791ccd87a · outbound
SR-Reward: Taking The Path More Traveled Revisiting successor features for inverse reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c12df2f-4ff4-44f1-a3ad-67a9b76dc95d · outbound
SR-Reward: Taking The Path More Traveled Deep inverse q-learning with constraints
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9a365a7-5379-48fd-83c1-7445ac51e84b · outbound
SR-Reward: Taking The Path More Traveled Imitation learning as f -divergence minimization, 2020
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51b6f6ec-49d9-46ed-afd8-cfba2e80545e · outbound
SR-Reward: Taking The Path More Traveled Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f40d1fbc-6f00-4ca6-9c75-cf8a8a4fd187 · outbound
SR-Reward: Taking The Path More Traveled Imitation learning via off-policy distribution matching
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d38b1494-317b-4fb6-9020-d3801ac98c92 · outbound
SR-Reward: Taking The Path More Traveled Offline reinforcement learning with implicit q-learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 492e82f3-6417-4ea7-8bec-4a96a47cb649 · outbound
SR-Reward: Taking The Path More Traveled Deep Successor Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baaab7a1-1292-4b63-995c-c19ef609b95e · outbound
SR-Reward: Taking The Path More Traveled Conservative q-learning for offline reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c855f70-6c92-4a1e-bf3a-4bbb17246f83 · outbound
SR-Reward: Taking The Path More Traveled Batch Reinforcement Learning, pp.\ 45--73
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfdfbf27-39bf-4cad-b1e2-d51bb413f8c2 · outbound
SR-Reward: Taking The Path More Traveled Energy-based imitation learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82b1b85c-db3f-4e4f-bcbc-7d1f4e203a34 · outbound
SR-Reward: Taking The Path More Traveled Learning self-correctable policies and value functions from demonstrations with negative sampling
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb73b68b-b885-4d72-b167-9fca4ae69a4b · outbound
SR-Reward: Taking The Path More Traveled Count-based exploration with the successor representation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c904107-37ba-4b24-875a-57b4b6dc007b · outbound
SR-Reward: Taking The Path More Traveled Rusu, Joel Veness, Marc G
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d98138-d117-444a-a2db-4837b87d4b78 · outbound
SR-Reward: Taking The Path More Traveled A first-occupancy representation for reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 385f08cf-0c36-4d1d-82dc-29d1e280304d · outbound
SR-Reward: Taking The Path More Traveled Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections, 2019
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed1cd8c7-c50a-4591-b519-1eaf0f81858e · outbound
SR-Reward: Taking The Path More Traveled Ng and Stuart Russell
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 920dbc8a-5536-4cb4-8ad1-5217836fdb9a · outbound
SR-Reward: Taking The Path More Traveled Bridging state and history representations: Understanding self-predictive rl, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 270bce38-e486-4197-8d02-1701a146b962 · outbound
SR-Reward: Taking The Path More Traveled Efficient training of artificial neural networks for autonomous navigation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc88be0-7288-4e31-8794-233f4ad09433 · outbound
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc410ba-2fc0-4486-ad18-38a88682a827 · outbound
SR-Reward: Taking The Path More Traveled Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638f87b3-4306-490b-a380-387fc0236d02 · outbound
SR-Reward: Taking The Path More Traveled Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 302cd67a-e909-4572-b325-0ddc319b4d27 · outbound
SR-Reward: Taking The Path More Traveled A motion retargeting method for effective mimicry-based teleoperation of robot arms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e0b5629e-d842-44e3-b7a8-11593c3bcab7 · outbound
SR-Reward: Taking The Path More Traveled Dragan, and Sergey Levine
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f7a51b9-4a93-49ee-bd76-ff0481931ac2 · outbound
SR-Reward: Taking The Path More Traveled A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6826ce05-3b2d-4427-a5dd-9fa7d11300d1 · outbound
SR-Reward: Taking The Path More Traveled Dual rl: Unification and new methods for reinforcement and imitation learning, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 313881f8-6826-4a4d-8e63-71e03ff540b7 · outbound
SR-Reward: Taking The Path More Traveled Lewis, and A
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b231494-24e3-4c0f-b569-5c8a1594e147 · outbound
SR-Reward: Taking The Path More Traveled Issues in using function approximation for reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3722a0c4-167e-44aa-b916-c7e4f949d9e0 · outbound
SR-Reward: Taking The Path More Traveled Mujoco: A physics engine for model-based control
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7423c8d-55a4-4073-ad02-9f5fe3e79d43 · outbound
SR-Reward: Taking The Path More Traveled Munchausen reinforcement learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20523608-9296-4a7f-8d1f-f9188dde7aa6 · outbound
SR-Reward: Taking The Path More Traveled Daydreamer: World models for physical robot learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c83658-6a74-4aee-9caf-8e85680b5d09 · outbound
SR-Reward: Taking The Path More Traveled Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators, 2023
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06728c8-6e80-444e-9aae-d8bba8838704 · outbound
SR-Reward: Taking The Path More Traveled Offline rl with no ood actions: In-sample learning via implicit value regularization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 054f9da7-b1eb-4f61-9bb6-18f39e83b3e3 · outbound
SR-Reward: Taking The Path More Traveled Deep reinforcement learning with successor features for navigation across similar environments
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 93c77db1-874a-465f-8d44-05beaddeadc4 · outbound
SR-Reward: Taking The Path More Traveled write newline
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.