Pith. sign in

REVIEW 2 cited by

Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.24128 v1 pith:FTPKCJFU submitted 2024-10-31 cs.LG

classification cs.LG
keywords mdpsquantilealgorithmdecompositionq-learningalgorithmsconvergenceperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents' preferences for certain outcomes. This paper proposes a new Q-learning algorithm for quantile optimization in MDPs with strong convergence and performance guarantees. The algorithm leverages a new, simple dynamic program (DP) decomposition for quantile MDPs. Compared with prior work, our DP decomposition requires neither known transition probabilities nor solving complex saddle point equations and serves as a suitable foundation for other model-free RL algorithms. Our numerical results in tabular domains show that our Q-learning algorithm converges to its DP variant and outperforms earlier algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Computing Monetary Risk Measures in Linear Time

    cs.LG 2026-07 accept novelty 6.0 of 10

    QuickVaR and QuickDivergence compute VaR and EWS φ-divergence risk measures (CVaR, TVaR) in expected O(n) time by avoiding full sorts via Quickselect-style partitioning and polymatroid structure.

  2. Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

    cs.LG 2026-02 conditional novelty 5.0 of 10

    A shifted-value transformation turns static CVaR MDPs into a bounded, contracting Bellman operator with dense rewards, enabling discretized value iteration and Q-learning with explicit error bounds.

Pith tools