Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:13.369371Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2502.03369.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:13.369371Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7d390698-79ce-46a3-a28b-e79ecaabfe62 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Agent-Agnostic Human-in-the-Loop Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f779148-5b50-4f46-ab15-ea9f0849d2da · outbound
Learning from Active Human Involvement through Proxy Value Propagation Constrained policy optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6551bf87-7e4c-4241-9215-45c03fb950f1 · outbound
Learning from Active Human Involvement through Proxy Value Propagation An interactive framework for learning continuous actions policies based on corrective feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f113462-bbb9-4ff1-bc2c-4d1475b62433 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Minimalistic gridworld environment for openai gym
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43d96413-8c44-46c5-a6ed-b9ad7354217f · outbound
Learning from Active Human Involvement through Proxy Value Propagation Christiano, Jan Leike, Tom B
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61c1ff7-38cd-4a67-a958-2eb43de93f67 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Open Problems in Cooperative AI
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef4784d-5dbf-4407-9428-04d38b84f334 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Magnetic control of tokamak plasmas through deep reinforcement learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5cbdc2e8-a9e3-4a6e-92db-05407621c93c · outbound
Learning from Active Human Involvement through Proxy Value Propagation CARLA: An open urban driving simulator
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5614c5d-7d61-4601-9830-9e1c162b2a6f · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning robust rewards with adverserial inverse reinforcement learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b39b107-02c1-44ed-804f-de83d3e30c30 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Addressing function approximation error in actor-critic methods
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7222c1b5-b6b1-4e9e-a42b-73c4e1fd9997 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8aac2172-d86a-40e0-b6c3-1092cba33595 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning to walk in the real world with minimal human effort, 2020
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5814ac2b-f61b-4abc-a940-19b105a23d71 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a12310a7-8404-448e-b892-16650f1dd590 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Generative adversarial imitation learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d9f4182-d0a1-46b2-a062-953718f16410 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Precise Synthetic Image and LiDAR (PreSIL) Dataset for Autonomous Vehicle Perception
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d88299e-4df6-4c4a-8b8b-cf1c3d882959 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning to share autonomy across repeated interaction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbbdf038-5ea7-40bf-b2db-ddc029307981 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Hg-dagger: Interactive imitation learning with human experts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 515442c5-2d78-450d-8dab-92b0fa12c494 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning to drive in a day
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12e15a05-11a0-42cc-a893-0fdd3527d49e · outbound
Learning from Active Human Involvement through Proxy Value Propagation Reinforcement learning from human reward: Discounting in episodic tasks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ec8c7b7-888c-4308-a845-7c7bf919bb68 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Specification gaming: the flip side of ai ingenuity
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eae9d699-d97b-4fa8-8a8f-b935d18a6caf · outbound
Learning from Active Human Involvement through Proxy Value Propagation Conservative q-learning for offline reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 480fa63f-cede-40c0-b88d-6a38f8e293db · outbound
Learning from Active Human Involvement through Proxy Value Propagation PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5c7713-fe20-4ef9-a9f2-ad7da4b506d9 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Scalable agent alignment via reward modeling: a research direction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800cc48d-9186-41e3-bec0-9b126f665946 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41784e27-ff14-4fe3-a9ec-8d65cb273d50 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 613635c1-2af8-4d10-9e08-ba39b93faf50 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Efficient learning of safe driving policy via human-ai copilot optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b5d3aa1-10ac-46a2-adf8-93d386449437 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Interactive learning from policy-dependent human feedback
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c17b998b-b5a3-4ca8-bef6-efdf4058b2bf · outbound
Learning from Active Human Involvement through Proxy Value Propagation Where to add actions in human-in- the-loop reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1010a489-38aa-4c5b-b1b4-c5e7c20ef80b · outbound
Learning from Active Human Involvement through Proxy Value Propagation Human-in-the-Loop Imitation Learning using Remote Teleoperation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60815919-ae78-415b-9ebc-0f7f148902dc · outbound
Learning from Active Human Involvement through Proxy Value Propagation Ensembledagger: A bayesian approach to safe imitation learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ced1ce6b-ef94-4b41-b8bb-c876ef40c9b6 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Human-level control through deep reinforcement learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3bba32-8576-4491-8fe4-a9c3e45548cb · outbound
Learning from Active Human Involvement through Proxy Value Propagation Interactively shaping robot behaviour with unlabeled human instructions
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 485f51ee-91ce-4237-afda-1ca95d0b9bde · outbound
Learning from Active Human Involvement through Proxy Value Propagation Deep exploration via bootstrapped DQN
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ee7e1128-c1dd-45f1-be1b-0f85cc47d448 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Training language models to follow instructions with human feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f37cc8a-19ae-4da6-b640-437947344033 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Deeptake: Prediction of driver takeover behavior using multimodal data
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d77a5421-8af5-435e-a7f1-2d5400274dd5 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning reward functions by integrating human demonstrations and preferences
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3504d69c-1fdf-4a9a-b3b2-fb81bd3064b2 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Stable-baselines3: Reliable reinforcement learning implementations
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4aa344c-5a33-4a04-b2c4-819b55859e9f · outbound
Learning from Active Human Involvement through Proxy Value Propagation Shared autonomy via deep reinforcement learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a7b02cd-c928-4ab8-8a6f-e684a891d10e · outbound
Learning from Active Human Involvement through Proxy Value Propagation Efficient reductions for imitation learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 512fd1f9-fd5d-493b-8cf5-ce44b21155d6 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Human compatible: Artificial intelligence and the problem of control
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef5d8521-5e86-45ba-8c00-014e4e319f1a · outbound
Learning from Active Human Involvement through Proxy Value Propagation Active preference-based learning of reward functions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43b7a1ab-eb22-4dd8-92f6-0e5c6ad147e6 · outbound
Learning from Active Human Involvement through Proxy Value Propagation The StarCraft Multi-Agent Challenge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b959a36c-801b-41d3-becb-2a8a2b0779bf · outbound
Learning from Active Human Involvement through Proxy Value Propagation Trial without error: Towards safe reinforcement learning via human intervention
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a52999c7-3485-4783-b7a0-55bf2061e842 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Prioritized Experience Replay
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d15502-5157-47ab-8d20-976a467ab046 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Proximal Policy Optimization Algorithms
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0398c6e5-90a9-4f6b-a123-049b82a417d6 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Mastering the game of go with deep neural networks and tree search
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd6d2e5-f549-4dab-a89d-eb87b501cb25 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Learning from interventions
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3b36b0b-6a65-4bec-bebb-593a2260a381 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Responsive safety in reinforcement learning by PID lagrangian methods
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16f063e0-5ee7-4fe0-a36f-97e2e770258a · outbound
Learning from Active Human Involvement through Proxy Value Propagation Intervention aided reinforcement learning for safe and practical policy optimization in navigation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f9894bd-0bf9-4556-8a26-6636f5946c53 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Appli: Adaptive planner parameter learning from interventions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 478fe29a-200e-4244-aaec-9c1429f8ed03 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Apple: Adaptive planner parameter learning from evaluative feedback
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6731b1a-369e-46d1-ab34-e27019e4ca7c · outbound
Learning from Active Human Involvement through Proxy Value Propagation Waytowich, Vernon Lawhern, and Peter Stone
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9efe55a1-6684-4c74-95cc-3541b6f622e2 · outbound
Learning from Active Human Involvement through Proxy Value Propagation A survey of preference-based reinforcement learning methods
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fff5f2f2-54f5-4012-8fb2-f5dd6ee30d87 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Look before you leap: Safe model-based reinforcement learning with human intervention
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a6b2247-a7b6-4c53-8a33-62eec05ac09a · outbound
Learning from Active Human Involvement through Proxy Value Propagation How to Leverage Unlabeled Data in Offline Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cba8519-29cb-44e7-beb5-2e893ce4d370 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Query-efficient imitation learning for end-to-end simulated driving
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e950529c-f767-45f5-9422-d19e5ac28403 · outbound
Learning from Active Human Involvement through Proxy Value Propagation We also compare the behavior of agents learned from PVP and TD3 baseline
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6068c2d7-ed81-48b2-89b1-a6aefef1393f · outbound
Learning from Active Human Involvement through Proxy Value Propagation We present the behavior comparison between PVP and TD3 baseline
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9bf4e8f1-30c6-4542-af09-7e41c28670bf · outbound
Learning from Active Human Involvement through Proxy Value Propagation PVP performs well in GTA V and can drive smoothly on the highway
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c5bcf62e-6cee-4520-b026-ec406bf02213 · outbound
Learning from Active Human Involvement through Proxy Value Propagation If the agent drives in the wrong way then the displace- ment reward will be multiplied by −1
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c52ab57e-f03e-4c80-a06c-8884e315b45d · outbound
Learning from Active Human Involvement through Proxy Value Propagation If the agent drives in wrong way then the speed reward will be multiplied by −1
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 17d20be8-c420-46e8-acdc-458c05267c24 · outbound
Learning from Active Human Involvement through Proxy Value Propagation Otherwise, it is 0
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0fc79af-f7b3-490a-81a4-4890185e7071 · outbound
Learning from Active Human Involvement through Proxy Value Propagation At that step, we set Rdisp = Rspeed = Rcollision = 0and assign Rterm according to the terminal state
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.