Pith. sign in

REVIEW 3 cited by

Reinforcement Learning for Jump-Diffusions, with Financial Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16449 v5 pith:XRFQUSC6 submitted 2024-05-26 cs.LG math.OCq-fin.MF

Reinforcement Learning for Jump-Diffusions, with Financial Applications

classification cs.LG math.OCq-fin.MF
keywords jump-diffusionlearningalgorithmscontroldiffusiondynamicsexploratorygeneral
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We study continuous-time reinforcement learning (RL) for stochastic control in which system dynamics are governed by jump-diffusion processes. We formulate an entropy-regularized exploratory control problem with stochastic policies to capture the exploration--exploitation balance essential for RL. Unlike the pure diffusion case initially studied by Wang et al. (2020), the derivation of the exploratory dynamics under jump-diffusions calls for a careful formulation of the jump part. Through a theoretical analysis, we find that one can simply use the same policy evaluation and $q$-learning algorithms in Jia and Zhou (2022a, 2023), originally developed for controlled diffusions, without needing to check a priori whether the underlying data come from a pure diffusion or a jump-diffusion. However, we show that the presence of jumps ought to affect parameterizations of actors and critics in general. We investigate as an application the mean--variance portfolio selection problem with stock price modelled as a jump-diffusion, and show that both RL algorithms and parameterizations are invariant with respect to jumps. Finally, we present a detailed study on applying the general theory to option hedging.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management

    cs.LG 2024-06 unverdicted novelty 7.0

    A continuous-time RL framework for intensity control in choice-based network revenue management outperforms discretization-based methods while scaling to large problems.

  2. Entropy Regularized Reinforcement Learning for Zero-Sum Stochastic Differential Games in a Regime-Switching Jump-Diffusion Process

    cs.LG 2026-06 unverdicted novelty 6.0

    Introduces ERRL-ZSSDGs with coupled HJBI equations from DPP, semi-analytical LQ solutions via ODEs, and Actor-Critic approximation for general regime-switching jump-diffusion cases, illustrated on an investment game.

  3. Continuous-time reinforcement learning for optimal switching over multiple regimes

    math.OC 2025-12 conditional novelty 5.0

    An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.