Pith. sign in

A Pontryagin Perspective on Reinforcement Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Reinforcement learning has traditionally focused on learning state-dependent policies to solve optimal control problems in a closed-loop fashion. In this work, we introduce the paradigm of open-loop reinforcement learning where a fixed action sequence is learned instead. We present three new algorithms: one robust model-based method and two sample-efficient model-free methods. Rather than basing our algorithms on Bellman's equation from dynamic programming, our work builds on Pontryagin's principle from the theory of open-loop optimal control. We provide convergence guarantees and evaluate all methods empirically on a pendulum swing-up task, as well as on two high-dimensional MuJoCo tasks, significantly outperforming existing baselines.

citation-role summary

background 1

citation-polarity summary

fields

eess.SY 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

DATA-DRIVEN PRONTO: a Model-free Solution for Numerical Optimal Control

eess.SY · 2025-06-18 · conditional · novelty 6.0

DATA-DRIVEN PRONTO iteratively estimates local linearized dynamics from perturbed closed-loop experiments, solves an LQR subproblem with those estimates, and provably converges to a neighborhood of the optimal solution whose size shrinks with the exploration dither amplitude.

citing papers explorer

Showing 1 of 1 citing paper.

  • DATA-DRIVEN PRONTO: a Model-free Solution for Numerical Optimal Control eess.SY · 2025-06-18 · conditional · none · ref 14 · internal anchor

    DATA-DRIVEN PRONTO iteratively estimates local linearized dynamics from perturbed closed-loop experiments, solves an LQR subproblem with those estimates, and provably converges to a neighborhood of the optimal solution whose size shrinks with the exploration dither amplitude.