In Webots-simulated maze navigation, Option-Critic HRL converged faster than PPO on two harder mazes, but the paper's evidence that sub-goals cause this advantage is compromised by confounded ablations.
Maximum a Posteriori Policy Optimisation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Hierarchical Reinforcement Learning in Multi-Goal Spatial Navigation with Autonomous Mobile Robots
In Webots-simulated maze navigation, Option-Critic HRL converged faster than PPO on two harder mazes, but the paper's evidence that sub-goals cause this advantage is compromised by confounded ablations.