Adding annealed Gaussian noise to rewards can help RL exploration, but this paper's proof of that claim is invalid and its SAC algorithm actually uses biased, non-zero-mean noise.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Exploration by Random Reward Perturbation
Adding annealed Gaussian noise to rewards can help RL exploration, but this paper's proof of that claim is invalid and its SAC algorithm actually uses biased, non-zero-mean noise.