Pith. sign in

REVIEW 2 cited by

Reward Estimation for Variance Reduction in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.03359 v2 pith:7MDUP5J6 submitted 2018-05-09 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords rewardlearningcorruptedhandlingrewardssignalstochasticvariance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning (RL) agents require the specification of a reward signal for learning behaviours. However, introduction of corrupt or stochastic rewards can yield high variance in learning. Such corruption may be a direct result of goal misspecification, randomness in the reward signal, or correlation of the reward with external factors that are not known to the agent. Corruption or stochasticity of the reward signal can be especially problematic in robotics, where goal specification can be particularly difficult for complex tasks. While many variance reduction techniques have been studied to improve the robustness of the RL process, handling such stochastic or corrupted reward structures remains difficult. As an alternative for handling this scenario in model-free RL methods, we suggest using an estimator for both rewards and value functions. We demonstrate that this improves performance under corrupted stochastic rewards in both the tabular and non-linear function approximation settings for a variety of noise types and environments. The use of reward estimation is a robust and easy-to-implement improvement for handling corrupted reward signals in model-free RL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hopfield Networks is All You Need

    cs.NE 2020-07 unverdicted novelty 7.0 of 10

    Modern Hopfield networks store exponentially many patterns, retrieve them in one update, and have an update rule equivalent to transformer attention, enabling new Hopfield layers that improve results on multiple insta...

  2. Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed

    cs.IR 2024-11 reject novelty 5.0 of 10

    SL-MGAC combines supervised reward prediction, user-group decomposition, and actor-critic RL to allocate live streams in a feed; offline and online tests report gains, but the reward predictor is partly fed the true r...

Pith tools