A hand-designed nested reward formula with a secondary multi-output network is claimed to improve PPO on Pendulum-v1, but the method is underspecified and the evidence is anecdotal.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Creating Hierarchical Dispositions of Needs in an Agent
A hand-designed nested reward formula with a secondary multi-output network is claimed to improve PPO on Pendulum-v1, but the method is underspecified and the evidence is anecdotal.