Pith. sign in

REVIEW 3 cited by

Deep Reinforcement Learning for Online Optimal Execution Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13493 v1 pith:AUOQFTAG submitted 2024-10-17 cs.LG stat.ML

Deep Reinforcement Learning for Online Optimal Execution Strategies

classification cs.LG stat.ML
keywords executionoptimalalgorithmlearningdecaydeepreinforcementstrategies
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper tackles the challenge of learning non-Markovian optimal execution strategies in dynamic financial markets. We introduce a novel actor-critic algorithm based on Deep Deterministic Policy Gradient (DDPG) to address this issue, with a focus on transient price impact modeled by a general decay kernel. Through numerical experiments with various decay kernels, we show that our algorithm successfully approximates the optimal execution strategy. Additionally, the proposed algorithm demonstrates adaptability to evolving market conditions, where parameters fluctuate over time. Our findings also show that modern reinforcement learning algorithms can provide a solution that reduces the need for frequent and inefficient human intervention in optimal execution tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can Reinforcement Learning Efficiently Discover Price Manipulation?

    q-fin.TR 2026-07 conditional novelty 6.0

    Under intermediate volatility and limited samples, model-free DDPG finds dynamic-arbitrage strategies more reliably than SLSQP run on noisily estimated Almgren-Chriss impact parameters, even though the latter knows th...

  2. Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

    q-fin.CP 2026-05 unverdicted novelty 5.0

    In a two-agent Almgren-Chriss liquidation game, deep RL agents given intra-episode history of prices and own actions achieve supra-competitive outcomes more frequently and persistently than agents without such memory.

  3. TT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution

    cs.AI 2026-06 unverdicted novelty 4.0

    TT-DAC-PS, an enhanced version of TD3, achieves lower mean implementation shortfall than PPO, SAC, A2C, TWAP, VWAP, and AC on LOB data from ten U.S. stocks.