Pith. sign in

REVIEW 2 cited by

Reinforcement Learning for Market Making in a Multi-agent Dealer Market

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.05892 v1 pith:TL3MDAOE submitted 2019-11-14 q-fin.TR cs.LGcs.MA

classification q-fin.TRcs.LGcs.MA
keywords marketagentinventorylearningmakerreinforcementdealerformulations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Market makers play an important role in providing liquidity to markets by continuously quoting prices at which they are willing to buy and sell, and managing inventory risk. In this paper, we build a multi-agent simulation of a dealer market and demonstrate that it can be used to understand the behavior of a reinforcement learning (RL) based market maker agent. We use the simulator to train an RL-based market maker agent with different competitive scenarios, reward formulations and market price trends (drifts). We show that the reinforcement learning agent is able to learn about its competitor's pricing policy; it also learns to manage inventory by smartly selecting asymmetric prices on the buy and sell sides (skewing), and maintaining a positive (or negative) inventory depending on whether the market price drift is positive (or negative). Finally, we propose and test reward formulations for creating risk averse RL-based market maker agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. Provably Convergent Actor-Critic for MARL through Risk-aversion

    cs.MA 2026-02 conditional novelty 6.0 of 10

    A fast-actor/slow-critic actor-critic algorithm converges with finite-sample guarantees to the unique risk-averse quantal-response equilibrium (RQE) of any general-sum Markov game, provided the agent-side parameters s...

  2. Learning Market Making with Closing Auctions

    q-fin.TR 2026-01 conditional novelty 5.0 of 10

    A neural-fitted Q-learning market maker that anticipates the closing auction beats Avellaneda-Stoikov and TWAP benchmarks on mean returns in the paper's simulations.

Pith tools