Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T05:24:53.844059Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2606.26527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T05:24:53.844059Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf04b379-a558-485b-9187-494c8668fe6d · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Deep reinforcement learning for autonomous driving: A survey,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e17af3-df6a-4c55-a763-0e71219471f5 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning for autonomous lane changing using set-based prediction,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce97bd5-f466-400d-af0c-9d9953b00104 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unsupervised reinforcement learning for multi-task autonomous driving: Expanding skills and cultivating curiosity,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c14c1ad-31b6-4e3d-92ee-058642c07588 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Driving tasks transfer using deep reinforcement learning for decision-making of autonomous vehicles in unsignalized intersection,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d9fff5-edaa-4ff9-9209-7d806b0a3493 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy A perspective of q-value estimation on offline-to-online reinforcement learning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d5d787-db8f-43ce-b0e2-7e1a84d496e7 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Sim-to-lab-to-real: Safe reinforcement learning with shielding and generalization guarantees,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591fc16c-fcbc-446b-93fd-438e2e3138f4 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Knowledge transfer from simple to complex: A safe and efficient reinforcement learning framework for autonomous driving decision-making,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59c97a9-9156-47a6-a5bd-e86894a34787 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Zero-shot deep reinforcement learning driving policy transfer for autonomous vehicles based on robust control,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51924c67-6b7c-436b-b71b-447134fbb78f · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safety reinforcement learning control via transfer learning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b113a94-afcf-4c4a-9f7a-6b229e3bceb3 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Federated Transfer Reinforcement Learning for Autonomous Driving
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6a73046-0631-4f66-b92c-e81d32f3f35a · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Scenario- level knowledge transfer for motion planning of autonomous driving via successor representation,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c83d9e7-d7dd-4c78-b75e-f76744b697c4 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Self-supervised domain transfer for reinforcement learning-based autonomous driving agent,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece1c202-224d-45b4-8661-ddc911f8f7ef · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Cross-domain adaptive transfer reinforcement learning based on state-action correspondence,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6cc2845-4143-4127-9651-c2d2b5564dfc · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 51573923-6671-471a-8f60-903a588abe1a · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Policy optimization with demonstrations,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7506a50e-5af8-4a9c-b942-0fa355e5c99c · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5a818391-6a13-46d5-baae-a766260f1c2a · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Knowledge transfer for deep reinforcement learning with hierarchical experience replay,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd322c82-940b-4847-9f50-009e2b09887b · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Improving reinforcement learning with confidence-based demonstrations,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac78be33-341c-4ed4-9482-7bc04e89ead8 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An enhanced advising model in teacher-student framework using state categorization,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35443eb2-b724-492f-a2a7-c1e5e229d202 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Human as ai mentor: En- hanced human-in-the-loop reinforcement learning for safe and efficient autonomous driving,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff7d50e-9183-4796-a818-4f3a7854372d · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Adaptive action advising with different rewards,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44423618-ab82-467b-9f54-4ef25c0745a1 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning via shielding,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cbaf01e-e927-48f0-89f0-76929bceba3f · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Safe reinforcement learning via shielding under partial observability,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9008c0-cc72-42e4-b01f-dee595077db0 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Robust model predictive shielding for safe reinforcement learning with stochastic dynamics,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091bdc8c-596e-4a05-96d4-1c8a5475428f · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Teaching on a budget in multi-agent deep reinforcement learning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09cf7116-3b99-46f5-b7f7-4c955abcde58 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Action advising with advice imitation in deep reinforcement learning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1819784a-bfcb-4a02-a127-2312f6c6190f · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Reinforcement learning with demonstrations from mismatched task under sparse reward,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e4480f-bf24-4c73-bc58-408d5628c2b3 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Psiphi- learning: Reinforcement learning with demonstrations using successor features and inverse temporal difference learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab690933-2014-4bb6-b2bb-e2e0fb4f15eb · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Hybrid reinforcement learning with expert state sequences,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1854cbe-9c52-4d5b-85a8-f8940ea56f1c · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Guided exploration with proximal policy optimization using a single demonstration,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251ee9eb-b1ee-4b4c-877d-25ae7f1b78be · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Hybrid rl: Using both offline and online data can make rl efficient,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4041bed7-64c6-4b9e-82db-c2f9f345862b · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2ca2ef-c774-4ebf-a232-e93879ef4e7f · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy DCUR: Data Curriculum for Teaching via Samples with Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ab869bf-1fba-4c0a-9ffc-792e8f05443d · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An actor-critic algorithm for constrained markov decision processes,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0578096c-f9a8-494a-928a-7a189c9dbeef · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Reinforcement learning by guided safe exploration,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2a8f051-75fd-4428-b5b5-5fb506f68ea7 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Guarded policy optimization with imperfect online demonstrations,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d22defa-47d6-4af8-a64f-17b48a3c876b · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Approximately optimal approximate rein- forcement learning,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87f193d-c0a9-4091-a0bc-7c65dd6697f1 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c18163e-bda3-44bd-9f4c-8dd0fd2778b3 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed8ab7f1-2817-4cbc-bfed-d933c3328404 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy An environment for autonomous driving decision-making,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a96287-ca84-4736-9699-37832c968bf5 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370964a7-e7ae-4569-ad97-1df63392fd99 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Congested traffic states in empirical observations and microscopic simulations,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15dea300-1ea8-436c-b2aa-4a247d2c1036 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Preferred time-headway of highway drivers,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84badfd9-f38a-4ea2-a1dd-947c2fb79cbd · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b27eee5-acbd-4cdb-a2e9-8d64c5f405d3 · outbound
Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy Responsive safety in reinforce- ment learning by pid lagrangian methods,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.