Pith. sign in

REVIEW 1 cited by

Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.04553 v2 pith:7TOOUNCG submitted 2025-05-07 q-fin.MF cs.AIq-fin.RM

classification q-fin.MFcs.AIq-fin.RM
keywords algorithmproposeauxiliaryclassconvexfunctionslearningproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

    cs.LG 2026-07 reject novelty 6.0 of 10

    A Bayesian IRL algorithm recovers an agent's distortion riskmetric from noisy binary choices at an exponential rate, while a PPO variant with a quantile network is proposed—though not proven—to optimize policies under...

Pith tools