Pith. sign in

REVIEW 2 cited by

RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.01972 v1 pith:AKCHYBNV submitted 2024-08-04 cs.LG

classification cs.LG
keywords rewardaveragecriteriontaskslearningmethodoff-policyreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we propose an off-policy deep reinforcement learning (DRL) method utilizing the average reward criterion. While most existing DRL methods employ the discounted reward criterion, this can potentially lead to a discrepancy between the training objective and performance metrics in continuing tasks, making the average reward criterion a recommended alternative. We introduce RVI-SAC, an extension of the state-of-the-art off-policy DRL method, Soft Actor-Critic (SAC), to the average reward criterion. Our proposal consists of (1) Critic updates based on RVI Q-learning, (2) Actor updates introduced by the average reward soft policy improvement theorem, and (3) automatic adjustment of Reset Cost enabling the average reward reinforcement learning to be applied to tasks with termination. We apply our method to the Gymnasium's Mujoco tasks, a subset of locomotion tasks, and demonstrate that RVI-SAC shows competitive performance compared to existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Empirical Study of Deep Reinforcement Learning in Continuing Tasks

    cs.AI 2025-01 conditional novelty 6.0 of 10

    An empirical study shows deep RL algorithms struggle in continuing tasks without resets and that TD-based reward centering improves their performance across larger MuJoCo and Atari testbeds.

  2. Average-Reward Soft Actor-Critic

    cs.LG 2025-01 reject novelty 4.0 of 10

    ASAC extends soft actor-critic to the entropy-regularized average-reward setting with a policy improvement theorem, but its claimed novelty is undermined by the earlier RVI-SAC algorithm.

Pith tools