Pith. sign in

REVIEW 2 cited by

Average-Reward Reinforcement Learning with Trust Region Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03442 v2 pith:BLJMMOKK submitted 2021-06-07 cs.LG cs.AI

classification cs.LGcs.AI
keywords averagediscountedcriterionlearningregionreinforcementtrustapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the discounted criterion is appropriate for certain tasks such as financial related problems, many engineering problems treat future rewards equally and prefer a long-run average criterion. In this paper, we study the reinforcement learning problem with the long-run average criterion. Firstly, we develop a unified trust region theory with discounted and average criteria and derive a novel performance bound within the trust region with the Perturbation Analysis (PA) theory. Secondly, we propose a practical algorithm named Average Policy Optimization (APO), which improves the value estimation with a novel technique named Average Value Constraint. Finally, experiments are conducted in the continuous control environment MuJoCo. In most tasks, APO performs better than the discounted PPO, which demonstrates the effectiveness of our approach. Our work provides a unified framework of the trust region approach including both the discounted and average criteria, which may complement the framework of reinforcement learning beyond the discounted objectives.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Empirical Study of Deep Reinforcement Learning in Continuing Tasks

    cs.AI 2025-01 conditional novelty 6.0 of 10

    An empirical study shows deep RL algorithms struggle in continuing tasks without resets and that TD-based reward centering improves their performance across larger MuJoCo and Atari testbeds.

  2. Average Reward Reinforcement Learning for Wireless Radio Resource Management

    cs.IT 2025-01 reject novelty 5.0 of 10

    Average reward RL, implemented as ARO-SAC, is reported to outperform discounted SAC by 15% in a RAN slicing radio resource management task.

Pith tools