Pith. sign in

REVIEW 3 cited by

Deep Bellman Hedging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00932 v4 pith:5C32FAFV submitted 2022-07-03 q-fin.CP q-fin.ST

classification q-fin.CPq-fin.ST
keywords hedgingapproachbellmandeepderivativesfinancialhedgeinstruments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an actor-critic-type reinforcement learning algorithm for solving the problem of hedging a portfolio of financial instruments such as securities and over-the-counter derivatives using purely historic data. The key characteristics of our approach are: the ability to hedge with derivatives such as forwards, swaps, futures, options; incorporation of trading frictions such as trading cost and liquidity constraints; applicability for any reasonable portfolio of financial instruments; realistic, continuous state and action spaces; and formal risk-adjusted return objectives. Most importantly, the trained model provides an optimal hedge for arbitrary initial portfolios and market states without the need for re-training. We also prove existence of finite solutions to our Bellman equation, and show the relation to our vanilla Deep Hedging approach

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Hedging Valuation Adjustment for Deep Hedging Policies under Market Frictions

    q-fin.RM 2026-07 conditional novelty 6.0 of 10

    A common-stress reserve framework for deep hedgers shows classical trading bands usually beat learned policies, with sparse learned execution winning only under a strict low-liquidity budget.

  2. Uncertainty-Aware Strategies: A Model-Agnostic Framework for Robust Financial Optimization through Subsampling

    q-fin.CP 2025-06 conditional novelty 5.0 of 10

    A framework that applies a risk measure on a distribution of models, approximated by subsampling, robustifies financial optimization against model uncertainty and scales via a memory-efficient CVaR-SGD algorithm.

  3. Risk-Averse Reinforcement Learning with Itakura-Saito Loss

    cs.LG 2025-05 conditional novelty 4.0 of 10

    The Itakura-Saito loss, derived from Bregman divergence, learns risk-averse value functions that match the exponential-utility Bellman equations and trains more stably than exponential MSE in the tested benchmarks.

Pith tools