Pith. sign in

REVIEW 1 cited by

Mean-Variance Efficient Reinforcement Learning with Applications to Dynamic Financial Investment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.01404 v4 pith:3FMKAPW2 submitted 2020-10-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords efficientpolicytrade-offapproachexpectedincreaselearningmean-variance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study investigates the mean-variance (MV) trade-off in reinforcement learning (RL), an instance of the sequential decision-making under uncertainty. Our objective is to obtain MV-efficient policies whose means and variances are located on the Pareto efficient frontier with respect to the MV trade-off; under the condition, any increase in the expected reward would necessitate a corresponding increase in variance, and vice versa. To this end, we propose a method that trains our policy to maximize the expected quadratic utility, defined as a weighted sum of the first and second moments of the rewards obtained through our policy. We subsequently demonstrate that the maximizer indeed qualifies as an MV-efficient policy. Previous studies that employed constrained optimization to address the MV trade-off have encountered computational challenges. However, our approach is more computationally efficient as it eliminates the need for gradient estimation of variance, a contributing factor to the double sampling issue observed in existing methodologies. Through experimentation, we validate the efficacy of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continuous-time optimal investment with portfolio constraints: a reinforcement learning approach

    q-fin.MF 2024-12 conditional novelty 5.0 of 10

    For entropy-regularized reinforcement learning in a continuous-time Merton market, the optimal exploratory policy is Gaussian, becoming truncated Gaussian under interval portfolio constraints, with closed-form log and...

Pith tools