REVIEW 3 cited by
Deep Bellman Hedging
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present an actor-critic-type reinforcement learning algorithm for solving the problem of hedging a portfolio of financial instruments such as securities and over-the-counter derivatives using purely historic data. The key characteristics of our approach are: the ability to hedge with derivatives such as forwards, swaps, futures, options; incorporation of trading frictions such as trading cost and liquidity constraints; applicability for any reasonable portfolio of financial instruments; realistic, continuous state and action spaces; and formal risk-adjusted return objectives. Most importantly, the trained model provides an optimal hedge for arbitrary initial portfolios and market states without the need for re-training. We also prove existence of finite solutions to our Bellman equation, and show the relation to our vanilla Deep Hedging approach
Forward citations
Cited by 3 Pith papers
-
Robust Hedging Valuation Adjustment for Deep Hedging Policies under Market Frictions
A common-stress reserve framework for deep hedgers shows classical trading bands usually beat learned policies, with sparse learned execution winning only under a strict low-liquidity budget.
-
Uncertainty-Aware Strategies: A Model-Agnostic Framework for Robust Financial Optimization through Subsampling
A framework that applies a risk measure on a distribution of models, approximated by subsampling, robustifies financial optimization against model uncertainty and scales via a memory-efficient CVaR-SGD algorithm.
-
Risk-Averse Reinforcement Learning with Itakura-Saito Loss
The Itakura-Saito loss, derived from Bregman divergence, learns risk-averse value functions that match the exponential-utility Bellman equations and trains more stably than exponential MSE in the tested benchmarks.
Discussion (0). Continue with ORCID to comment.