Pith. sign in

REVIEW 1 cited by

Deep Reinforcement Learning for Active High Frequency Trading

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.07107 v3 pith:GHDYCMSM submitted 2021-01-18 cs.LG cs.AIcs.MAq-fin.TR

classification cs.LGcs.AIcs.MAq-fin.TR
keywords dataagentstradingfrequencyhightrainingableactive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the Proximal Policy Optimization algorithm. The training is performed on three contiguous months of high frequency Limit Order Book data, of which the last month constitutes the validation data. In order to maximise the signal to noise ratio in the training data, we compose the latter by only selecting training samples with largest price changes. The test is then carried out on the following month of data. Hyperparameters are tuned using the Sequential Model Based Optimization technique. We consider three different state characterizations, which differ in their LOB-based meta-features. Analysing the agents' performances on test data, we argue that the agents are able to create a dynamic representation of the underlying environment. They identify occasional regularities present in the data and exploit them to create long-term profitable trading strategies. Indeed, agents learn trading strategies able to produce stable positive returns in spite of the highly stochastic and non-stationary environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents

    cs.LG 2025-07 conditional novelty 5.0 of 10

    EXPLAIN trains an interpretable linear policy from an expert's offline trajectories by combining advantage-weighted policy gradients with a behavioral cloning regularizer.

Pith tools