BIOMAP achieves the optimal reward on the Mask Cliff Walking benchmark by reconstructing the state graph from action vectors, but this hinges on an unstated assumption that states are geometric positions.
Planning and control in stochastic domains with imperfect information
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.SY 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
A Model-free Biomimetics Algorithm for Deterministic Partially Observable Markov Decision Process
BIOMAP achieves the optimal reward on the Mask Cliff Walking benchmark by reconstructing the state graph from action vectors, but this hinges on an unstated assumption that states are geometric positions.