WOMBET transfers experience by generating filtered, uncertainty-penalized rollouts from a source world model and adaptively mixing them with target online data for sample-efficient RL.
Cog: Connecting new skills to past experience with offline reinforcement learning.arXiv preprint arXiv:2010.14500
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2verdicts
UNVERDICTED 2representative citing papers
VIPO improves model-based offline RL by minimizing value function inconsistency between direct data estimates and model predictions, achieving SOTA results on D4RL and NeoRL benchmarks.
citing papers explorer
-
WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning
WOMBET transfers experience by generating filtered, uncertainty-penalized rollouts from a source world model and adaptively mixing them with target online data for sample-efficient RL.
-
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
VIPO improves model-based offline RL by minimizing value function inconsistency between direct data estimates and model predictions, achieving SOTA results on D4RL and NeoRL benchmarks.