Pith. sign in

REVIEW 1 cited by

Quantile-Based Deep Reinforcement Learning using Two-Timescale Policy Gradient Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07248 v1 pith:NY3K4MEJ submitted 2023-05-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords policyalgorithmsquantilequantile-basedalgorithmcumulativedeepgradient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classical reinforcement learning (RL) aims to optimize the expected cumulative reward. In this work, we consider the RL setting where the goal is to optimize the quantile of the cumulative reward. We parameterize the policy controlling actions by neural networks, and propose a novel policy gradient algorithm called Quantile-Based Policy Optimization (QPO) and its variant Quantile-Based Proximal Policy Optimization (QPPO) for solving deep RL problems with quantile objectives. QPO uses two coupled iterations running at different timescales for simultaneously updating quantiles and policy parameters, whereas QPPO is an off-policy version of QPO that allows multiple updates of parameters during one simulation episode, leading to improved algorithm efficiency. Our numerical results indicate that the proposed algorithms outperform the existing baseline algorithms under the quantile criterion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborating in a competitive world: Heterogeneous Multi-Agent Decision Making in Symbiotic Supply Chain Environments

    cs.MA 2025-01 conditional novelty 5.0 of 10

    Separate per-node policies reduce the bullwhip effect in a simulated two-node supply chain, but a single shared policy earns more in low-demand settings; SAC beats PPO in high demand.

Pith tools