Pith. sign in

REVIEW 2 cited by

Parametric Return Density Estimation for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1203.3497 v1 pith:WSDIUX5U submitted 2012-03-15 cs.LG stat.ML

classification cs.LGstat.ML
keywords densityalgorithmsexpectedparametricreturnreturnsconditionalcriteria
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most conventional Reinforcement Learning (RL) algorithms aim to optimize decision-making rules in terms of the expected returns. However, especially for risk management purposes, other risk-sensitive criteria such as the value-at-risk or the expected shortfall are sometimes preferred in real applications. Here, we describe a parametric method for estimating density of the returns, which allows us to handle various criteria in a unified manner. We first extend the Bellman equation for the conditional expected return to cover a conditional probability density of the returns. Then we derive an extension of the TD-learning algorithm for estimating the return densities in an unknown environment. As test instances, several parametric density estimation algorithms are presented for the Gaussian, Laplace, and skewed Laplace distributions. We show that these algorithms lead to risk-sensitive as well as robust RL paradigms through numerical experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Practical Risk Measures in Reinforcement Learning

    cs.LG 2019-08 reject novelty 5.0 of 10

    An actor-critic algorithm with a Monte Carlo risk critic is proposed for optimizing reinforcement learning policies under arbitrary, possibly non-coherent risk measures, with a risk function fitted from simulated data.

  2. Bridging Econometrics and AI: VaR Estimation via Reinforcement Learning and GARCH Models

    cs.AI 2025-04 conditional novelty 4.0 of 10

    A Double Deep Q-Network that classifies days as low- or high-risk is used to scale GARCH-based Value-at-Risk, reducing violations and capital requirements on daily Euro Stoxx 50 data from 2008 to 2025.

Pith tools