SimuDICE uses DualDICE weights and model confidence to re-sample synthetic transitions in a tabular world model, improving offline Dyna-Q style policy optimization in small discrete environments.
In: Proceedings of the 36th International Conference on Ma- chine Learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
SimuDICE uses DualDICE weights and model confidence to re-sample synthetic transitions in a tabular world model, improving offline Dyna-Q style policy optimization in small discrete environments.