A model-based RL framework that alternates causal structure learning with empowerment-driven exploration, plus a curiosity reward, improves sample efficiency and asymptotic performance in six environments.
Moreover, The policyπcollect is trained with a reward function r = tanh(PdS j=1 log p(sj t+1|st,at) p(sj t+1|PAsj ) )
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Empowerment Gain through Causal Structure Learning in Model-Based RL
A model-based RL framework that alternates causal structure learning with empowerment-driven exploration, plus a curiosity reward, improves sample efficiency and asymptotic performance in six environments.