REVIEW 2 cited by
Reward Centering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards' empirical average. The improvement is substantial at commonly used discount factors and increases further as the discount factor approaches one. In addition, we show that if a problem's rewards are shifted by a constant, then standard methods perform much worse, whereas methods with reward centering are unaffected. Estimating the average reward is straightforward in the on-policy setting; we propose a slightly more sophisticated method for the off-policy setting. Reward centering is a general idea, so we expect almost every reinforcement-learning algorithm to benefit by the addition of reward centering.
Forward citations
Cited by 2 Pith papers
-
Harnessing the Power of Reinforcement Learning for Adaptive MCMC
A contrastive-divergence-based reward stabilizes reinforcement learning of position-dependent step sizes for gradient-based MCMC, beating constant-step tuning on most of 44 benchmark posteriors.
-
Learning from Expert Factors: Trajectory-level Reward Shaping for Formulaic Alpha Mining
A trajectory-level reward shaping method for RL-based formulaic alpha mining uses exact subsequence matching against expert formulas and reward centering to accelerate training and slightly improve mined factors.
Discussion (0). Sign in to comment.