Pith. sign in

REVIEW 1 cited by

Robust Market Making via Adversarial Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.01820 v2 pith:ODGLASTC submitted 2020-03-03 q-fin.TR cs.AIcs.LGstat.ML

classification q-fin.TRcs.AIcs.LGstat.ML
keywords marketadversarialadversaryagentsempiricallygamelearningmaker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that adversarial reinforcement learning (ARL) can be used to produce market marking agents that are robust to adversarial and adaptively-chosen market conditions. To apply ARL, we turn the well-studied single-agent model of Avellaneda and Stoikov [2008] into a discrete-time zero-sum game between a market maker and adversary. The adversary acts as a proxy for other market participants that would like to profit at the market maker's expense. We empirically compare two conventional single-agent RL agents with ARL, and show that our ARL approach leads to: 1) the emergence of risk-averse behaviour without constraints or domain-specific penalties; 2) significant improvements in performance across a set of standard metrics, evaluated with or without an adversary in the test environment, and; 3) improved robustness to model uncertainty. We empirically demonstrate that our ARL method consistently converges, and we prove for several special cases that the profiles that we converge to correspond to Nash equilibria in a simplified single-stage game.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility

    q-fin.TR 2025-08 conditional novelty 4.0 of 10

    A 4-action reinforcement learning market maker trained with Hawkes order arrivals at low volatility continues to provide two-sided quotes over 92% of the time and holds stable Sharpe ratios when tested at 100x higher ...

Pith tools