REVIEW 2 cited by
Discounted Reinforcement Learning Is Not an Optimization Problem
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks. It is not an optimization problem in its usual formulation, so when using function approximation there is no optimal policy. We substantiate these claims, then go on to address some misconceptions about discounting and its connection to the average reward formulation. We encourage researchers to adopt rigorous optimization approaches, such as maximizing average reward, for reinforcement learning in continuing tasks.
Forward citations
Cited by 2 Pith papers
-
EVAL: EigenVector-based Average-reward Learning
EVAL learns the optimal policy for entropy-regularized average-reward MDPs by training neural networks to approximate the dominant eigenvector of a tilted transition matrix, with a variant that recovers the unregulari...
-
Average-Reward Soft Actor-Critic
ASAC extends soft actor-critic to the entropy-regularized average-reward setting with a policy improvement theorem, but its claimed novelty is undermined by the earlier RVI-SAC algorithm.
Discussion (0). Continue with ORCID to comment.