Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:10:38.484582Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2512.17091.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:10:38.484582Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:53:31.123181Z
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 89493903-9a99-473b-bd31-69f08f0dc9be · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making OpenAI Gym
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2a29e33-de26-42e6-82cb-2e8a6753f5a4 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making UCB Exploration via Q-Ensembles
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 889510dd-7412-4fe0-8376-98e4ce3c289a · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Td-mpc2: Scalable, robust world models for continuous control
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e391eb19-83b1-46e9-95a4-ca85572e563b · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b080c14-6b01-42eb-ae29-1bbe6dd8a330 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Real-time gait adaptation for quadrupeds using model predictive control and reinforcement learning.arXiv preprint arXiv:2510.20706
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b59c09ec-9c11-48b0-96f7-51ad64493e79 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Unifying Model Predictive Path Integral Control, Reinforcement Learning, and Diffusion Models for Optimal Control and Planning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b2d108c-c455-4b83-96be-81a7c3c95c00 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e95f15e8-a5c2-4aef-8e82-d9b827f48e9a · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Moerland, Joost Broekens, Aske Plaat, and Catholijn M
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f65740bf-e880-4694-a923-f7fcc0acd06a · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9266b73-c226-402c-a52d-e0b4edcc0c09 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Actor-critic model predictive control: Differentiable optimization meets reinforcement learn- ing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 17906c10-c23e-47c4-b177-effce9049956 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a2b69ba-8b4f-4e69-b893-e00b80a8a1c1 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e26ba27-c60b-4ac5-abba-bb997a5fcd8a · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Yubin Wang, Zengqi Peng, Yusen Xie, Yulin Li, Hakim Ghazzai, and Jun Ma
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a910ddc-52fe-47f8-b80d-5b345c1eaa9d · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Learning the References of Online Model Predictive Control for Urban Self-Driving
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9971c900-d870-435b-ac2e-c08eb1f4ef07 · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making A KL-regularization Framework for Learning to Plan with Adaptive Priors
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07131240-f5e1-40ea-97c3-c2165f0b1a9c · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making • Passing reward: increased with the vehicle’s relative speed to a nearbyOtherVehiclewhen overtaking (i.e., larger forward relative velocity yields larger reward)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3563ce6c-d570-4e30-bf70-805a17e327da · outbound
Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making RL term.Three common forms are used for the RL term
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10bc8b39-f458-4823-b588-dfd021e47d3c · inbound
Deep Reinforcement-Learning-Guided Model Predictive Control for Preventing Overtakes in Autonomous Racing Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.