Pith. sign in

Efficient Reinforcement Learning in Probabilistic Reward Machines

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this paper, we study reinforcement learning in Markov Decision Processes with Probabilistic Reward Machines (PRMs), a form of non-Markovian reward commonly found in robotics tasks. We design an algorithm for PRMs that achieves a regret bound of $\widetilde{O}(\sqrt{HOAT} + H^2O^2A^{3/2} + H\sqrt{T})$, where $H$ is the time horizon, $O$ is the number of observations, $A$ is the number of actions, and $T$ is the number of time-steps. This result improves over the best-known bound, $\widetilde{O}(H\sqrt{OAT})$ of \citet{pmlr-v206-bourel23a} for MDPs with Deterministic Reward Machines (DRMs), a special case of PRMs. When $T \geq H^3O^3A^2$ and $OA \geq H$, our regret bound leads to a regret of $\widetilde{O}(\sqrt{HOAT})$, which matches the established lower bound of $\Omega(\sqrt{HOAT})$ for MDPs with DRMs up to a logarithmic factor. To the best of our knowledge, this is the first efficient algorithm for PRMs. Additionally, we present a new simulation lemma for non-Markovian rewards, which enables reward-free exploration for any non-Markovian reward given access to an approximate planner. Complementing our theoretical findings, we show through extensive experiment evaluations that our algorithm indeed outperforms prior methods in various PRM environments.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Physics-Informed Reward Machines

cs.LG · 2025-08-14 · conditional · novelty 6.0

Physics-informed reward machines augment reward machines with ODE-driven continuous state, and experiments show that counterfactual experiences and reward shaping built on them speed up reinforcement learning.

citing papers explorer

Showing 1 of 1 citing paper.

  • Physics-Informed Reward Machines cs.LG · 2025-08-14 · conditional · none · ref 2006 · internal anchor

    Physics-informed reward machines augment reward machines with ODE-driven continuous state, and experiments show that counterfactual experiences and reward shaping built on them speed up reinforcement learning.