Pith. sign in

Paper Citation Record · LEDGER

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making

As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2512.17091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.17091 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T21:10:38.484582Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:53:31.123181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact11
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89493903-9a99-473b-bd31-69f08f0dc9be · outbound

This paper cites OpenAI Gym.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making OpenAI Gym

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:11:16.778701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:17afd1c6c1e1f8e637cc404875017691402b4407d15aeeda596ce836a26dd649

Observation c2a29e33-de26-42e6-82cb-2e8a6753f5a4 · outbound

This paper cites UCB Exploration via Q-Ensembles.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making UCB Exploration via Q-Ensembles

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:11:16.782360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:c7f41ec8fbc5c94a605a1d992aeac499ab1095d6ea079f5dcc2ffcb9cd4cf04b

Observation 889510dd-7412-4fe0-8376-98e4ce3c289a · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Td-mpc2: Scalable, robust world models for continuous control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:11:17.233074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:e19d06e4eef1a692615f55aeaea23853881281a42589a44c866680a670844ba5

Observation e391eb19-83b1-46e9-95a4-ca85572e563b · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:11:16.814065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:d67c9cb340d2e361ec7dd3643c3f98c03ccbcbdb9c3268f751fa80f8af361f58

Observation 7b080c14-6b01-42eb-ae29-1bbe6dd8a330 · outbound

This paper cites Real-time gait adaptation for quadrupeds using model predictive control and reinforcement learning.arXiv preprint arXiv:2510.20706.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Real-time gait adaptation for quadrupeds using model predictive control and reinforcement learning.arXiv preprint arXiv:2510.20706

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.795150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:420c04fb2038adbcaead64a529fdab50b26bcbb3347b79eaed941e571747cac7

Observation b59c09ec-9c11-48b0-96f7-51ad64493e79 · outbound

This paper cites Unifying Model Predictive Path Integral Control, Reinforcement Learning, and Diffusion Models for Optimal Control and Planning.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Unifying Model Predictive Path Integral Control, Reinforcement Learning, and Diffusion Models for Optimal Control and Planning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.806116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:cba9ead5542009bec25ab9cc8650169004d39fe15b72cb05909e8b7e505e1fe1

Observation 9b2d108c-c455-4b83-96be-81a7c3c95c00 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:11:16.810107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:87f5127ddd7d0665289d5dac32874febc5060622f1295b6c64b9c8d14ad81aa5

Observation e95f15e8-a5c2-4aef-8e82-d9b827f48e9a · outbound

This paper cites Moerland, Joost Broekens, Aske Plaat, and Catholijn M.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Moerland, Joost Broekens, Aske Plaat, and Catholijn M

Reference 8

Resolution
metadata mismatch
doi, observed 2026-05-16T21:11:16.357775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:a5a33728b6884e6764ea7042b17c9dda806d5d46ae2e6298b5ec0a38d93ac3d6

Observation f65740bf-e880-4694-a923-f7fcc0acd06a · outbound

This paper cites TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.790943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:54aaeeee2cf152b8f0cec004f2ff9f17b93277c268a631bed792f678e87ac03a

Observation d9266b73-c226-402c-a52d-e0b4edcc0c09 · outbound

This paper cites Actor-critic model predictive control: Differentiable optimization meets reinforcement learn- ing.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Actor-critic model predictive control: Differentiable optimization meets reinforcement learn- ing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.819087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:5c8642d78a1ef38889eb36d773743ba6905e7309c8bd254e59c4b5061948ed01

Observation 17906c10-c23e-47c4-b177-effce9049956 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:11:16.822808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:b6dd1151de9a2486024e223f1cc3c5c4b2943a7e0117e6f1754b4becc38873f3

Observation 1a2b69ba-8b4f-4e69-b893-e00b80a8a1c1 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:11:16.802600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:2faa4048a9842c827c0d2e0cb107525c52c7176ddf47a6417a6d8335d8abc35e

Observation 5e26ba27-c60b-4ac5-abba-bb997a5fcd8a · outbound

This paper cites Yubin Wang, Zengqi Peng, Yusen Xie, Yulin Li, Hakim Ghazzai, and Jun Ma.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Yubin Wang, Zengqi Peng, Yusen Xie, Yulin Li, Hakim Ghazzai, and Jun Ma

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.353694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:cfe70d0029aa3940a0c91d9268904ba1e271bd3cc7715117aad4ef4ec956af80

Observation 2a910ddc-52fe-47f8-b80d-5b345c1eaa9d · outbound

This paper cites Learning the References of Online Model Predictive Control for Urban Self-Driving.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making Learning the References of Online Model Predictive Control for Urban Self-Driving

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:11:16.799416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:40dd6fd5ff239d3b76145b5b18016fc5c6f7c5d2d74600286c1de981f50e7a30

Observation 9971c900-d870-435b-ac2e-c08eb1f4ef07 · outbound

This paper cites A KL-regularization Framework for Learning to Plan with Adaptive Priors.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making A KL-regularization Framework for Learning to Plan with Adaptive Priors

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:27.082124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:5e25f451368598f3b6bfbde8ce1d30b1ebed05c1a3bc325409cd4c3acff4d15d

Observation 07131240-f5e1-40ea-97c3-c2165f0b1a9c · outbound

This paper cites • Passing reward: increased with the vehicle’s relative speed to a nearbyOtherVehiclewhen overtaking (i.e., larger forward relative velocity yields larger reward).

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making • Passing reward: increased with the vehicle’s relative speed to a nearbyOtherVehiclewhen overtaking (i.e., larger forward relative velocity yields larger reward)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:11:17.238714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:efefa16d5559fc0db26a2de6848cf7525c6f932c32399ec9af6e83e3b09e2920

Observation 3563ce6c-d570-4e30-bf70-805a17e327da · outbound

This paper cites RL term.Three common forms are used for the RL term.

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making RL term.Three common forms are used for the RL term

Reference 17

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T21:11:17.235724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:10:38.484582Z digest=sha256:26ad2871249fb650dd1a20414ec4693d3a7988fbc5f847a0a9bb55843e830b14

Pith citing papers

Observation 10bc8b39-f458-4823-b588-dfd021e47d3c · inbound

Deep Reinforcement-Learning-Guided Model Predictive Control for Preventing Overtakes in Autonomous Racing cites this paper.

Deep Reinforcement-Learning-Guided Model Predictive Control for Preventing Overtakes in Autonomous Racing Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:53:31.123181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:53:31.123181Z digest=sha256:a5a07815c18f14163c409b9c967a78ae00a12954963b45ae6feff477776ab16f