Pith. sign in

Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution,

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2024 1

verdicts

REJECT 1

representative citing papers

Creating Hierarchical Dispositions of Needs in an Agent

cs.LG · 2024-11-23 · reject · novelty 3.0

A hand-designed nested reward formula with a secondary multi-output network is claimed to improve PPO on Pendulum-v1, but the method is underspecified and the evidence is anecdotal.

citing papers explorer

Showing 1 of 1 citing paper.

  • Creating Hierarchical Dispositions of Needs in an Agent cs.LG · 2024-11-23 · reject · none · ref 8

    A hand-designed nested reward formula with a secondary multi-output network is claimed to improve PPO on Pendulum-v1, but the method is underspecified and the evidence is anecdotal.