Pith. sign in

REVIEW 2 cited by

Reward Centering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09999 v2 pith:DLBMWF53 submitted 2024-05-16 cs.LG cs.AI

classification cs.LGcs.AI
keywords rewardcenteringmethodsrewardsadditionaveragediscountperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards' empirical average. The improvement is substantial at commonly used discount factors and increases further as the discount factor approaches one. In addition, we show that if a problem's rewards are shifted by a constant, then standard methods perform much worse, whereas methods with reward centering are unaffected. Estimating the average reward is straightforward in the on-policy setting; we propose a slightly more sophisticated method for the off-policy setting. Reward centering is a general idea, so we expect almost every reinforcement-learning algorithm to benefit by the addition of reward centering.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harnessing the Power of Reinforcement Learning for Adaptive MCMC

    stat.CO 2025-07 conditional novelty 7.0 of 10

    A contrastive-divergence-based reward stabilizes reinforcement learning of position-dependent step sizes for gradient-based MCMC, beating constant-step tuning on most of 44 benchmark posteriors.

  2. Learning from Expert Factors: Trajectory-level Reward Shaping for Formulaic Alpha Mining

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A trajectory-level reward shaping method for RL-based formulaic alpha mining uses exact subsequence matching against expert formulas and reward centering to accelerate training and slightly improve mined factors.

Pith tools