REVIEW 3 cited by
Recent Advances in Reinforcement Learning in Finance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid changes in the finance industry due to the increasing amount of data have revolutionized the techniques on data processing and data analysis and brought new theoretical and computational challenges. In contrast to classical stochastic control theory and other analytical approaches for solving financial decision-making problems that heavily reply on model assumptions, new developments from reinforcement learning (RL) are able to make full use of the large amount of financial data with fewer model assumptions and to improve decisions in complex financial environments. This survey paper aims to review the recent developments and use of RL approaches in finance. We give an introduction to Markov decision processes, which is the setting for many of the commonly used RL approaches. Various algorithms are then introduced with a focus on value and policy based methods that do not require any model assumptions. Connections are made with neural networks to extend the framework to encompass deep RL algorithms. Our survey concludes by discussing the application of these RL algorithms in a variety of decision-making problems in finance, including optimal execution, portfolio optimization, option pricing and hedging, market making, smart order routing, and robo-advising.
Forward citations
Cited by 3 Pith papers
-
Reinforcement-Learning Portfolio Allocation with Dynamic Embedding of Market Information
DERL, a dynamic-embedding RL framework, beats value/equal-weighted portfolios and an MLP predict-then-optimize baseline on 1993-2022 U.S. large-cap returns, chiefly in high-volatility periods.
-
Differentially Private Policy Gradient
A differentially private policy gradient method that frames DP clipping as a trust-region choice and demonstrates competitive returns on deep RL and RLHF tasks, with the caveat that harder tasks use weak privacy budgets.
-
Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk
VaR-CPO approximates non-differentiable VaR constraints via Cantelli's inequality to enable safe, sample-efficient policy optimization with zero training violations in feasible environments.
Discussion (0). Continue with ORCID to comment.