SensorOpt formulates backup sensor selection for RL policies as a budget-constrained QUBO using a second-order return approximation, and finds near-optimal configurations with Tabu Search.
Dynamic programming with incomplete information to overcome navigational uncertainty in a nautical environment
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Using a novel toy nautical navigation environment, we show that dynamic programming can be used when only incomplete information about a partially observed Markov decision process (POMDP) is known. By incorporating uncertainty into our model, we show that navigation policies can be constructed that maintain safety, outperforming the baseline performance of traditional dynamic programming for Markov decision processes (MDPs). Adding in controlled sensing methods, we show that these policies can also lower measurement costs at the same time.
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimizing Sensor Redundancy in Sequential Decision-Making Problems
SensorOpt formulates backup sensor selection for RL policies as a budget-constrained QUBO using a second-order return approximation, and finds near-optimal configurations with Tabu Search.