Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:47.517274Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.03568.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:47.517274Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b953d159-f33d-4b89-859a-3537b2c5e1ab · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to drive in a day,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea909278-58b2-43ac-9010-e526a626c62c · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End to End Learning for Self-Driving Cars
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b284c1-6c32-4c92-bd57-e8d5f8d7c391 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Dense reinforcement learning for safety validation of autonomous vehicles,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9599d1f-fd6d-4551-a9b6-ad054c3bd309 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey of deep RL and IL for autonomous driving policy learning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a10452b6-967f-4279-b900-5e0790312f10 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on autonomous vehicle control in the era of mixed-autonomy: From physics-based to ai-guided driving policy learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6c01a77-4c2a-4f80-91d0-67b85257cbb5 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep learning for safe autonomous driving: Current chal- lenges and future directions,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e94c2b-9b6d-496f-9258-84821eaebf43 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Survey of deep reinforcement learning for motion planning of autonomous vehicles,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33fc5d1f-f2a8-4f43-b19f-8a3c79b377b0 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbac7cfb-67d7-4cd9-9adb-34ae7100f6ce · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning for autonomous driving: A survey,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a382cf7-ee65-43e4-a623-303b4ec4c4fa · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving End-to-end urban driving by imitating a reinforcement learning coach,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69b0d92e-7f70-4794-a84a-fc54ae364c6a · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Reward misdesign for autonomous driving,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d2bb129-4ecb-4225-92d3-85b487ae0178 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5789285-6345-42fd-8fa0-d698171dc58a · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Demonstrating specification gaming in reasoning models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ecdb816-979a-4115-8595-9a98158cc2fd · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Toward human-in-the-loop AI: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e200d637-9a49-4c9b-bcf8-f7d3425a8012 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hindsight credit assignment,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37afddd2-4f33-41ff-8d6a-2707ab0e72bc · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7435c22e-26b8-487a-b920-9a468755d7a3 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A survey on imitation learning techniques for end-to-end autonomous vehicles,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43b0205f-a0b0-4a5c-84ba-28463da293fb · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conditional predictive behavior planning with inverse reinforcement learning for human-like autonomous driving,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52d91390-1107-4988-9158-bdb23e1d15f6 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Pattern recognition and adaptive control,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4696d27-913a-475b-a557-d457f26f83c7 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving An Algorithmic Perspective on Imitation Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f971b72-1e49-456d-8b9a-ff91feb8d2ef · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning a decision module by imitating driver’s control behaviors,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33693b59-16e0-45ff-8b2c-7ff49c725d71 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Safety Critics for Exploration
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab2d0dd-d2b1-4721-b75a-7bfb2264ad49 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavior Regularized Offline Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e335325-1715-4ac0-a391-11f2cc86de03 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Off-policy deep reinforcement learning without exploration,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0098b27e-2e4b-42a2-a36d-324d9d0fd4e7 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Adversarial inverse rein- forcement learning with self-attention dynamics model,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9434015b-8421-4144-b37f-18da759de9d6 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient reductions for imitation learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4442ef43-f0c1-432b-9059-456c2821fa2c · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Exploring the limi- tations of behavior cloning for autonomous driving,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bf27769-625a-48d1-ab28-945048db412f · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Behavioral cloning a correction,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec8175b5-24a6-42d4-bdcc-c042ea77d078 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 024cbaec-61db-4023-908b-e53548811411 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A reduction of imitation learning and structured prediction to no-regret online learning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f1763acf-d4a0-4831-a456-b69e1506e204 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Query-Efficient Imitation Learning for End-to-End Autonomous Driving
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70cfac0c-e7e9-4dfc-a259-3032037c6108 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Hg-dagger: Interactive imitation learning with human experts,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0258af-51c5-4dba-a66c-2a63b26941c0 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a20d808-d509-43e5-b925-2926ac9b8c0d · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Expert intervention learning,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93b1d758-f991-4c18-92d1-157a51a2a22b · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human-in-the-Loop Imitation Learning using Remote Teleoperation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc54a30-2cf5-4e27-bca1-c6b64468a705 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Deep reinforcement learning from human preferences,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 558d2604-d8a2-45cd-96b4-1858083f53f1 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Batch active preference-based learning of reward functions,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64f9f942-7044-4d1b-b85d-3446d11595d6 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning Reward Functions by Integrating Human Demonstrations and Preferences
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e09e0c-5b81-46f2-8a1d-49f877c648b0 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Efficient learning of safe driving policy via human-AI copilot optimization,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67936181-90d0-495a-866a-28f330be8d8c · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning from active human involvement through proxy value propagation,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ace12819-7be9-4dbb-8251-53a121e283b2 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 355d23c3-b0f1-4007-80e5-e2862f3e8c3b · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Socially situated artificial intelligence enables learning from human interaction,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 268b589d-5665-465a-8013-6846a2e03e11 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09d7c262-872f-4a76-b8bb-ab18800494a0 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Human as AI mentor: Enhanced human-in-the-loop reinforcement learning for safe and effi- cient autonomous driving,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd045378-b146-4e55-9cf1-b5559197a214 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Guarded policy optimization with imperfect online demonstrations,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d56724d3-6b63-4773-84a8-f20f074ef99b · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Trust region policy optimization,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbc2e929-024c-40ff-8841-12f2dfc434be · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b46c8c8a-366e-447a-bee1-d91116c5d1d4 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Responsive safety in reinforce- ment learning by pid lagrangian methods,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d55fd2c-4562-4271-8430-f18d391150b2 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Learning to Walk in the Real World with Minimal Human Effort
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f771110b-c8bf-4dda-8076-11d34b2454a5 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Conservative Q- learning for offline reinforcement learning,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b598292-45f2-4e28-9d94-55dfe16f77f7 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Proximal Policy Optimization Algorithms
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389de93b-200b-4f6b-aa22-ce32ef49726e · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b2a3f9b-c5c4-4ba3-8bb6-24d538ca433f · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving A framework for behavioural cloning,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee91778d-321a-480a-a9dd-e402f4b66a43 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Generative adversarial imitation learning,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec87f5b9-e218-4ec9-9066-043b308dffc2 · outbound
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.