Pith. sign in

REVIEW

Factors of Influence of the Overestimation Bias of Q-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05262 v1 pith:VJIDAVF2 submitted 2022-10-11 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords overestimationbiasinfluenceq-learningalgorithmalphagammaaccurate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study whether the learning rate $\alpha$, the discount factor $\gamma$ and the reward signal $r$ have an influence on the overestimation bias of the Q-Learning algorithm. Our preliminary results in environments which are stochastic and that require the use of neural networks as function approximators, show that all three parameters influence overestimation significantly. By carefully tuning $\alpha$ and $\gamma$, and by using an exponential moving average of $r$ in Q-Learning's temporal difference target, we show that the algorithm can learn value estimates that are more accurate than the ones of several other popular model-free methods that have addressed its overestimation bias in the past.

Discussion (0). Sign in to comment.

Pith tools