Umbrella RL adds an ensemble-entropy bonus to policy gradient to solve sparse-reward, trap-heavy RL tasks, and reports large gains over PPO, RND, iLQR, and value iteration on two toy benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems
Umbrella RL adds an ensemble-entropy bonus to policy gradient to solve sparse-reward, trap-heavy RL tasks, and reports large gains over PPO, RND, iLQR, and value iteration on two toy benchmarks.